⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Docker & Containers Interview Questions Scenario 158 of 158 in Docker & Containers
Staff Infrastructure Architect Docker Container Runtime & Systems Engineering Production Scenario

Q: A catastrophic production incident occurred on an active database node: the root filesystem reached 100% capacity. SREs ran `du -sh /var/lib/docker` which reported only 18GB used, yet `df -h` showed 180GB used (a 162GB discrepancy!). Running `docker system prune` reclaimed zero space. Meanwhile, containers began failing with I/O errors and databases entered read-only mode. You must author a postmortem root-cause analysis, explain the Linux filesystem discrepancy, and create an automated SRE operational triage runbook.

Conduct a root cause investigation for runaway disk space exhaustion in `/var/lib/docker`. Triage unlinked open file descriptors holding disk space (`lsof +L1`), container log bloat, and write amplification.

#Docker #Linux #SRE #Postmortem #Storage
🎙️ Candidate Opening & Architectural Context
"Conduct a root cause investigation for runaway disk space exhaustion in `/var/lib/docker`. Triage unlinked open file descriptors holding disk space (`lsof +L1`), container log bloat, and write amplification."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Docker Certified Associate (DCA) Hands-On Lab Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

Step 1

Explain the du vs df Discrepancy: Unlinked Open Files

Understand Linux VFS mechanics: `du` (disk usage) traverses directory trees and calculates file sizes based on visible filenames. `df` (disk free) queries the filesystem superblock for allocated disk blocks. When a file is deleted (`rm`) while an active process still holds its file descriptor open, its directory entry is removed (invisible to `du`), but the kernel cannot free the disk blocks until the process closes the file descriptor or terminates.

<!-- du vs df Discrepancy Mechanics -->
Process (Log Writer) ───[ Holds open File Descriptor (FD) ]───> Inode #918237 (160GB)
                                                                       │
Administrator executes: rm /var/lib/docker/containers/.../app.log     │
                                                                       │
Result:
- Directory link removed -> du cannot see it! (du reports 18GB)
- Inode blocks NOT freed -> df reports 180GB allocated! (100% Full!)
Pro Tip: Explain the du vs df Discrepancy: Unlinked Open Files
Step 2

Locate Unlinked Open File Descriptors Using lsof

Run `lsof +L1` to find open file descriptors on deleted files holding massive disk capacity.

# Find unlinked open files (> 1GB) holding disk space
sudo lsof +L1 /var/lib/docker | grep -E 'deleted|CONTAINER' | sort -nk 7 -r | head -n 10
# Sample output:
# COMMAND   PID USER   FD   TYPE DEVICE   SIZE/OFF NLINK   NODE NAME
# java     4120 root    4w   REG  259,1 171798691840     0 918237 /var/lib/docker/.../app.log (deleted)
Pro Tip: Locate Unlinked Open File Descriptors Using lsof
Advertisement
Step 3

Safely Release Disk Blocks Without Killing Mission-Critical Processes

To immediately free the disk blocks without restarting or terminating the running container, truncate the file descriptor directly via `/proc//fd/`.

# Truncate unlinked file descriptor to 0 bytes
sudo truncate -s 0 /proc/4120/fd/4

# Verify disk space is instantly reclaimed in df
df -h /var/lib/docker
# Output: Available: 162GB (Discrepancy resolved instantly!)
Pro Tip: Safely Release Disk Blocks Without Killing Mission-Critical Processes
Step 4

Implement Automated SRE Runbook and Monitoring Safeguards

Deploy an automated Prometheus alert tracking the delta between `node_filesystem_size_bytes - node_filesystem_free_bytes` and `du` scans. Add automated log truncation and logrotate policies.

# Prometheus Alert Rule
# expr: (node_filesystem_size_bytes{mountpoint="/var/lib/docker"} - node_filesystem_free_bytes{mountpoint="/var/lib/docker"}) / node_filesystem_size_bytes{mountpoint="/var/lib/docker"} > 0.85
# for: 5m
# labels:
#   severity: critical
# annotations:
#   summary: "Runaway Docker disk consumption detected! Check unlinked open files via lsof +L1."
Pro Tip: Implement Automated SRE Runbook and Monitoring Safeguards
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Discrepancies between `df` and `du` on Docker hosts are almost always caused by unlinked open files where logs were deleted while a process held the file descriptor open. Truncating `/proc/<pid>/fd/<n>` instantly recovers disk blocks without terminating the workload."
⚡ 60-Second Elevator Pitch Talking Points
  • W
  • e
  • d
  • i
  • a
  • g
  • n
  • o
  • s
  • e
  • d
  • a
  • s
  • e
  • v
  • e
  • r
  • e
  • 1
  • 6
  • 0
  • G
  • B
  • d
  • i
  • s
  • k
  • e
  • x
  • h
  • a
  • u
  • s
  • t
  • i
  • o
  • n
  • o
  • u
  • t
  • a
  • g
  • e
  • w
  • h
  • e
  • r
  • e
  • `
  • d
  • u
  • `
  • a
  • n
  • d
  • `
  • d
  • f
  • `
  • r
  • e
  • p
  • o
  • r
  • t
  • e
  • d
  • c
  • o
  • n
  • f
  • l
  • i
  • c
  • t
  • i
  • n
  • g
  • m
  • e
  • t
  • r
  • i
  • c
  • s
  • .
  • U
  • s
  • i
  • n
  • g
  • `
  • l
  • s
  • o
  • f
  • +
  • L
  • 1
  • `
  • ,
  • w
  • e
  • i
  • d
  • e
  • n
  • t
  • i
  • f
  • i
  • e
  • d
  • t
  • h
  • a
  • t
  • a
  • d
  • e
  • l
  • e
  • t
  • e
  • d
  • a
  • p
  • p
  • l
  • i
  • c
  • a
  • t
  • i
  • o
  • n
  • l
  • o
  • g
  • w
  • a
  • s
  • s
  • t
  • i
  • l
  • l
  • h
  • e
  • l
  • d
  • o
  • p
  • e
  • n
  • b
  • y
  • a
  • r
  • u
  • n
  • n
  • i
  • n
  • g
  • J
  • a
  • v
  • a
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • ,
  • p
  • r
  • e
  • v
  • e
  • n
  • t
  • i
  • n
  • g
  • t
  • h
  • e
  • k
  • e
  • r
  • n
  • e
  • l
  • f
  • r
  • o
  • m
  • f
  • r
  • e
  • e
  • i
  • n
  • g
  • d
  • i
  • s
  • k
  • b
  • l
  • o
  • c
  • k
  • s
  • .
  • B
  • y
  • t
  • r
  • u
  • n
  • c
  • a
  • t
  • i
  • n
  • g
  • t
  • h
  • e
  • f
  • i
  • l
  • e
  • d
  • e
  • s
  • c
  • r
  • i
  • p
  • t
  • o
  • r
  • d
  • i
  • r
  • e
  • c
  • t
  • l
  • y
  • t
  • h
  • r
  • o
  • u
  • g
  • h
  • `
  • /
  • p
  • r
  • o
  • c
  • /
  • <
  • p
  • i
  • d
  • >
  • /
  • f
  • d
  • /
  • `
  • ,
  • w
  • e
  • i
  • n
  • s
  • t
  • a
  • n
  • t
  • l
  • y
  • f
  • r
  • e
  • e
  • d
  • 1
  • 6
  • 0
  • G
  • B
  • o
  • f
  • d
  • i
  • s
  • k
  • s
  • p
  • a
  • c
  • e
  • w
  • i
  • t
  • h
  • o
  • u
  • t
  • b
  • o
  • u
  • n
  • c
  • i
  • n
  • g
  • t
  • h
  • e
  • a
  • p
  • p
  • l
  • i
  • c
  • a
  • t
  • i
  • o
  • n
  • .
Advertisement
Want more Docker & Containers scenarios?
Explore our complete collection of scenario-based Docker & Containers interview runbooks.
Browse All Docker & Containers Questions →