Q: A server's disks are incredibly slow. `iostat -x 1` shows %util at 100%. How do you identify whether it's latency on reading data or writing data?
While %util shows the disk is completely saturated (100% time spent doing I/O), it doesn't explain the load profile.
#Linux #Linux / SRE — Scenario-Based Interview Questions #L2 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""We encountered this OS-level bottleneck during peak traffic and diagnosed it down to kernel and filesystem metrics. The interviewer is testing: Deciphering complex I/O metrics, `iostat` columns.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
While %util shows the disk is completely saturated (100% time spent doing I/O), it doesn't explain the load profile.
r_awaitvsw_await: These denote the average wait time (in milliseconds) for Read requests vs Write requests to be served. Ifr_awaitis 5ms andw_awaitis 500ms, the disk is severely struggling to write.rkB/svswkB/s: These denote the actual volume of Kilobytes Read vs Written per second.
2️⃣
Remediation & Permanent Safeguards
In the iostat output, I would focus on two specific column groups: If the write latency is massive, turning on Write-Back caching on a RAID controller, or migrating to an SSD, is the immediate hardware solution.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: r_await vs w_await: These denote the average wait time (in milliseconds) for Read requests vs Write requests to be served. If r_aw."
⚡ 60-Second Elevator Pitch Talking Points
- r_await vs w_await: These denote the average wait time (in milliseconds) for Read requests vs Wri...
- rkB/s vs wkB/s: These denote the actual volume of Kilobytes Read vs Written per second.
Advertisement