Skip to main content

Troubleshooting · Linux and System Administration

Deleted Files but Disk Still Full: Diagnose df versus du

Deleting a pathname does not close open file descriptors. Storage can remain allocated after the last link is removed while a process holds the file open. df measures filesystem…

Technical review:

Architecture and operating model

Deleting a pathname does not close open file descriptors. Storage can remain allocated after the last link is removed while a process holds the file open. df measures filesystem allocation and du reachable named content, so they can differ.

First verify the mount and inode utilization. lsof +L1 finds open unlinked files but permissions/namespaces can limit visibility. Snapshots, reserved blocks and files hidden beneath another mount can also cause differences.

Do not kill production processes randomly or blindly truncate /proc/PID/fd paths. Close descriptors using supported log-reopen/reload or controlled restart procedures. Test log rotation to prevent recurrence.

  1. 1df/du discrepancy
  2. 2Mount/inode check
  3. 3Open unlinked descriptor
  4. 4Controlled reopen/retest
Verify the mount point.

Design parameters

Mount boundary
Use du -x for one filesystem and account for container mount namespaces.
Descriptor owner
Record PID, service, size and link count together.
Inode
Inode exhaustion can block writes despite free GB; inspect df -i separately.

Worked example

A lab process holds a 5 GB log open. Removing its pathname can reduce du by 5 GB while df stays unchanged. A supported reopen closes the descriptor and releases space. Validate the same behavior during real log rotation.

Example commands: replace lab values and confirm permissions and software versions before use.

df -hT
df -i
du -xsh /var/log
sudo lsof +L1
# Use the owning application’s supported reopen/reload procedure after review.

Troubleshooting

ObservationLikely cause / distinctionVerification
du small, df largeOpen deleted file or snapshot.Inspect lsof +L1 and filesystem snapshot use.
Free GB but writes failInode or quota exhaustion.Check df -i and user quota.

Acceptance checks

  1. Verify the mount point.
  2. Measure inode usage.
  3. Find the descriptor owner.
  4. Use application-supported reopen.
  5. Remeasure df/du.
  6. Retest log rotation.

Related concepts

Filesystem and storage layers

An application directory, mount point, logical volume and physical device belong to different layers. Inode exhaustion, read-only mounts and filesystem errors can stop writes even when space appears available. A snapshot often shares the same storage failure domain and is not an independent backup. Mount options and expansion methods depend on the filesystem. Verify mounts after reboot: writing to an unmounted directory can fill the wrong device.

Service lifecycle and dependencies

A running process does not prove that users receive correct service. Review startup order, network or database dependencies, environment variables, file access and health checks together. Restart loops can conceal the root cause. Compare exit codes, recent logs and resource limits, and identify whether systemd or container management triggered the restart. After a controlled restart, validate sessions, writes and dependent services.

Capacity and usable headroom

Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.

Telemetry and time correlation

Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.

File permissions and service identities

File access combines users/groups, permission bits, ACLs and security policies. Running an application as root can hide permission problems; prefer a justified, restricted service identity in production. Directory traversal permission differs from file-read permission. Sharing permissions do not necessarily override filesystem restrictions. Identify the actual runtime user, directory chain and ACLs when troubleshooting. Test narrowly required access rather than opening permissions globally.

Availability versus recovery

High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.

Primary documentation

Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.

Knowledge Center

Enterprise IT Product Sales, Licensing and Deployment
Enterprise IT Project & Solution Scenarios
View all related content
Text on WhatsApp
Copied!