⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Linux Linux / SRE — Scenario-Based Interview Questions Staff SRE Scenario [L3]

Q: A database server shows 100% disk utilization on iostat, but random read latency is acceptable while sequential write latency spikes wildly. How do you distinguish between read bottlenecks and write bottlenecks, and what tools reveal the root cause?

Disk latency has two distinct bottlenecks: read latency (affecting query performance) and write latency (affecting commits/fsync). They r...

#Linux #Linux / SRE — Scenario-Based Interview Questions #L3 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain my systematic Linux troubleshooting methodology using Brendan Gregg's USE method. The interviewer is testing: Advanced I/O profiling, disk subsystem diagnostics, queue depth.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Disk latency has two distinct bottlenecks: read latency (affecting query performance) and write latency (affecting commits/fsync). They require different tools to diagnose.

  • r/s and w/s: Read and write operations per second (independently measured).
  • rareq-sz and wareq-sz: Average request sizes (small vs large).
  • r_await and w_await: Average latency for reads and writes (in ms).
  • svctm: Service time (how long the disk takes to serve one request).
  • %util: Percentage of time the disk is busy (100% = fully saturated).
2️⃣

Remediation & Permanent Safeguards

Tool 1: iostat -x 1 shows: If r_await is 2ms but w_await is 500ms, writes are the bottleneck, not reads. Tool 2: blktrace + blkparse captures every disk I/O operation: Output shows timestamp, operation type (R/W), block address, size, queue depth, and latency for each I/O. You can see if writes are getting stuck behind a large read or being reordered by the scheduler. Root causes:

blktrace -d /dev/sda -o - | blkparse -i -
  • High write latency + queue depth > 1: Disk controller or SSD firmware throttling writes (e.g., NAND garbage collection).
  • High write latency + queue depth = 1: Single slow write (e.g., journal flush to slow media or encryption overhead).
  • Random writes slower than sequential: RAID write-back cache eviction or lack of RAID optimization for random workloads.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: r/s and w/s: Read and write operations per second (independently measured).."
⚡ 60-Second Elevator Pitch Talking Points
  • r/s and w/s: Read and write operations per second (independently measured).
  • rareq-sz and wareq-sz: Average request sizes (small vs large).
  • r_await and w_await: Average latency for reads and writes (in ms).
Advertisement
Want more Linux scenarios?
Explore our complete collection of scenario-based Linux interview runbooks.
Browse All Linux Questions →

📚 Related Production Scenarios in Linux