⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Linux Performance & SRE Core SRE Methodology

Q: Production server becomes slow — how do you identify CPU/memory/disk/network issues?

Mastering the USE Method (Utilization, Saturation, Errors) to isolate system bottlenecks across CPU, Memory, Disk I/O, and Network in under 60 seconds.

#Linux #USE Method #vmstat #iostat #sar #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When a production server becomes slow, I apply Brendan Gregg's USE Method (Utilization, Saturation, Errors) using a systematic 60-second command checklist to isolate which subsystem is bottlenecking."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

First 15 Seconds: System Overview

Assess overall system pressure:

  • uptime: Look at load average (1m, 5m, 15m). Compare against core count (nproc). If 4 cores and load average is 28, system is heavily saturated.
  • dmesg -T | tail -30: Check kernel ring buffer for hardware errors, TCP drop warnings, or OOM killer events.
2️⃣

CPU vs Disk I/O: vmstat 1 5

The most informative single command in Linux:

  • Run: vmstat 1 5 and read the columns:
  • r (run queue): Processes waiting for CPU. If r > total cores, CPU is saturated.
  • b (blocked): Processes blocked waiting for I/O. If b > 0, system is waiting on disk/NFS.
  • us (user) vs sy (system): High us = application code. High sy = kernel context switching or excessive syscalls.
  • wa (iowait): High wa (>20%) = CPU is idle because it is waiting for disk/storage to respond! The bottleneck is DISK, not CPU!
3️⃣

Memory Check: free -h & vmstat swap

Inspect real available memory and paging:

  • free -h: Look at available (not free). Available accounts for reclaimable page cache and buffers.
  • In vmstat 1 5, look at si (swap in) and so (swap out). If so > 0, the system is actively thrashing memory to disk, causing severe latency.
  • dmesg -T | grep -i 'killed process': Check if kernel OOM killer has executed.
4️⃣

Disk I/O & Network Deep Dive

Isolate storage or packet bottlenecks:

  • Disk I/O (iostat -xz 1 5): Check %util and await. If %util > 90% or await > 15ms, physical EBS/disk storage is saturated.
  • Network (sar -n DEV 1 5): Check rxpck/s, txpck/s, and rxdrop/txdrop for dropped packets.
  • Socket Pressure (ss -s): Check TCP socket counts. If timewait or estab is in tens of thousands, connection pool exhaustion is occurring.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Use the USE Method. Run 'uptime' -> 'vmstat 1 5' -> 'free -h' -> 'iostat -xz 1 5'. High 'wa' in vmstat means Disk I/O bottleneck; high 'si/so' means memory swapping; high 'r' means CPU saturation."
⚡ 60-Second Elevator Pitch Talking Points
  • Step 1: 'uptime' -> compare 1/5/15m load average against 'nproc' core count.
  • Step 2: 'vmstat 1 5' -> check 'r' (CPU saturation), 'b' (I/O blocked), and 'wa' (if iowait is high, bottleneck is disk, not CPU).
  • Step 3: 'free -h' -> check 'available' memory; check 'si/so' in vmstat for swap thrashing.
  • Step 4: 'iostat -xz 1 5' -> check '%util' and 'await' response times for storage bottlenecks.
  • Step 5: 'sar -n DEV 1 5' and 'ss -s' -> check network throughput, dropped packets, and socket saturation.
  • Step 6: 'dmesg -T | tail -30' for OOM kills and hardware errors.
Advertisement
Want more Linux scenarios?
Explore our complete collection of scenario-based Linux interview runbooks.
Browse All Linux Questions →

📚 Related Production Scenarios in Linux