Q: Disk reaches 100% — how would you find and safely remove the cause?
Production runbook for recovering from 100% disk utilization: inode exhaustion, du vs df discrepancies, the unlinked open file descriptor trap (lsof deleted), and safe log truncation.
#Linux #df #du #lsof #Deleted Files Trap #Inodes #Storage
🎙️ Candidate Opening & Architectural Context
"When a disk hits 100%, services begin dropping connections, crash, and cannot write logs or create PID files. The investigation requires checking both disk blocks and inodes, and avoiding the classic 'deleted file still held open' trap."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Check Mounts and Inodes (df -h & df -i)
Identify which filesystem is full and whether it is bytes or inodes:
df -h: Identify which partition is at 100% (e.g./,/var, or/tmp).- The Inode Trap (
df -i): Ifdf -hshows 40% free space but applications report 'No space left on device', you have run out of Inodes! Millions of tiny 0-byte files (like session files in/var/spoolor PHP sessions) exhaust inodes even when gigabytes of disk remain.
2️⃣
Locate Largest Directories and Files (du)
Scan the filesystem safely without crossing mount boundaries:
du -ahx / | sort -rh | head -20: Lists top 20 space-consuming files/directories.- Why
-xflag is critical: The-xflag stays on one filesystem. Without it,duscans NFS mounts, Docker volumes, and virtual filesystems (/proc), crashing or taking hours. - Quick check on common culprits:
du -sh /var/log/* /var/lib/docker/* | sort -h.
3️⃣
The 'Deleted File Held Open' Trap (lsof +L1)
The #1 issue senior SREs look for:
- An engineer deletes a 40GB log file using
rm /var/log/app.log. Butdf -hSTILL shows 100% full! Why? - Linux Filesystem Mechanics: If a running process (NGINX, Java, database) has the file open, deleting the file removes the directory directory entry, but the OS cannot free the disk blocks until the file descriptor is closed!
- Command to identify:
lsof +L1orlsof | grep deleted. - How to fix without restarting service: Truncate the open file descriptor directly via
/proc:> /proc/<PID>/fd/<FD_NUMBER>(instantly frees the disk blocks without restarting the app!).
4️⃣
Safe Cleanup Best Practices
How to free space immediately without breaking services:
- Never rm an active log file: Always truncate:
truncate -s 0 /var/log/app.logor> /var/log/app.log. - Vacuum systemd journal logs:
journalctl --vacuum-time=2dorjournalctl --vacuum-size=500M. - Clean Docker leftovers:
docker system prune -af --volumes(verify unused volumes first). - Permanent fix: Configure
logrotatewith compression and retention limits.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Check 'df -i' for inode exhaustion. Use 'lsof +L1' to find deleted files held open by running processes, and truncate them via /proc/<PID>/fd/<FD> rather than rebooting."
⚡ 60-Second Elevator Pitch Talking Points
- Check blocks & inodes: 'df -h' for disk space, 'df -i' for inode exhaustion (millions of tiny files).
- Find large files: 'du -ahx / | sort -rh | head -20' (use -x to avoid traversing other mounts/NFS).
- Check deleted files held open: 'lsof +L1' or 'lsof | grep deleted' for deleted files still locked by active processes.
- Fix open deleted files: Truncate via '> /proc/<PID>/fd/<FD>' to free disk space immediately without restarting process.
- Clean safely: Truncate logs (truncate -s 0 file.log) rather than 'rm'; run 'journalctl --vacuum-size=500M'.
- Prevent: Set up logrotate, CloudWatch disk space alarm at 80%, and automated cleanup jobs.
Advertisement