Troubleshooting · Linux and System Administration
Deleted Files but Disk Still Full: Diagnose df versus du
Deleting a pathname does not close open file descriptors. Storage can remain allocated after the last link is removed while a process holds the file open. df measures filesystem…
Technical review:
Architecture and operating model
Deleting a pathname does not close open file descriptors. Storage can remain allocated after the last link is removed while a process holds the file open. df measures filesystem allocation and du reachable named content, so they can differ.
First verify the mount and inode utilization. lsof +L1 finds open unlinked files but permissions/namespaces can limit visibility. Snapshots, reserved blocks and files hidden beneath another mount can also cause differences.
Do not kill production processes randomly or blindly truncate /proc/PID/fd paths. Close descriptors using supported log-reopen/reload or controlled restart procedures. Test log rotation to prevent recurrence.
- 1df/du discrepancy
- 2Mount/inode check
- 3Open unlinked descriptor
- 4Controlled reopen/retest
Design parameters
- Mount boundary
- Use du -x for one filesystem and account for container mount namespaces.
- Descriptor owner
- Record PID, service, size and link count together.
- Inode
- Inode exhaustion can block writes despite free GB; inspect df -i separately.
Worked example
A lab process holds a 5 GB log open. Removing its pathname can reduce du by 5 GB while df stays unchanged. A supported reopen closes the descriptor and releases space. Validate the same behavior during real log rotation.
Example commands: replace lab values and confirm permissions and software versions before use.
df -hT
df -i
du -xsh /var/log
sudo lsof +L1
# Use the owning application’s supported reopen/reload procedure after review.
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| du small, df large | Open deleted file or snapshot. | Inspect lsof +L1 and filesystem snapshot use. |
| Free GB but writes fail | Inode or quota exhaustion. | Check df -i and user quota. |
Acceptance checks
- Verify the mount point.
- Measure inode usage.
- Find the descriptor owner.
- Use application-supported reopen.
- Remeasure df/du.
- Retest log rotation.
Related concepts
Filesystem and storage layers
An application directory, mount point, logical volume and physical device belong to different layers. Inode exhaustion, read-only mounts and filesystem errors can stop writes even when space appears available. A snapshot often shares the same storage failure domain and is not an independent backup. Mount options and expansion methods depend on the filesystem. Verify mounts after reboot: writing to an unmounted directory can fill the wrong device.
Service lifecycle and dependencies
A running process does not prove that users receive correct service. Review startup order, network or database dependencies, environment variables, file access and health checks together. Restart loops can conceal the root cause. Compare exit codes, recent logs and resource limits, and identify whether systemd or container management triggered the restart. After a controlled restart, validate sessions, writes and dependent services.
Capacity and usable headroom
Raw capacity is not the capacity available to applications. RAID or erasure coding, filesystems, reserved space, metadata, snapshots and growth headroom are separate deductions. TB and TiB representations also change the displayed number. Write calculations with units, establish protected usable capacity, then subtract operating reserves. Track growth rate as well as current utilization. The projected exhaustion date should leave enough time to procure and deploy additional capacity.
Telemetry and time correlation
Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.
File permissions and service identities
File access combines users/groups, permission bits, ACLs and security policies. Running an application as root can hide permission problems; prefer a justified, restricted service identity in production. Directory traversal permission differs from file-read permission. Sharing permissions do not necessarily override filesystem restrictions. Identify the actual runtime user, directory chain and ACLs when troubleshooting. Test narrowly required access rather than opening permissions globally.
Availability versus recovery
High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.