Q: Production server becomes slow — how do you identify CPU/memory/disk/network issues?
Mastering the USE Method (Utilization, Saturation, Errors) to isolate system bottlenecks across CPU, Memory, Disk I/O, and Network in under 60 seconds.
#Linux #USE Method #vmstat #iostat #sar #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When a production server becomes slow, I apply Brendan Gregg's USE Method (Utilization, Saturation, Errors) using a systematic 60-second command checklist to isolate which subsystem is bottlenecking."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
First 15 Seconds: System Overview
Assess overall system pressure:
uptime: Look at load average (1m, 5m, 15m). Compare against core count (nproc). If 4 cores and load average is 28, system is heavily saturated.dmesg -T | tail -30: Check kernel ring buffer for hardware errors, TCP drop warnings, or OOM killer events.
2️⃣
CPU vs Disk I/O: vmstat 1 5
The most informative single command in Linux:
- Run:
vmstat 1 5and read the columns: - r (run queue): Processes waiting for CPU. If r > total cores, CPU is saturated.
- b (blocked): Processes blocked waiting for I/O. If b > 0, system is waiting on disk/NFS.
- us (user) vs sy (system): High
us= application code. Highsy= kernel context switching or excessive syscalls. - wa (iowait): High
wa(>20%) = CPU is idle because it is waiting for disk/storage to respond! The bottleneck is DISK, not CPU!
3️⃣
Memory Check: free -h & vmstat swap
Inspect real available memory and paging:
free -h: Look at available (not free). Available accounts for reclaimable page cache and buffers.- In
vmstat 1 5, look at si (swap in) and so (swap out). Ifso > 0, the system is actively thrashing memory to disk, causing severe latency. dmesg -T | grep -i 'killed process': Check if kernel OOM killer has executed.
4️⃣
Disk I/O & Network Deep Dive
Isolate storage or packet bottlenecks:
- Disk I/O (
iostat -xz 1 5): Check%utilandawait. If%util > 90%orawait > 15ms, physical EBS/disk storage is saturated. - Network (
sar -n DEV 1 5): Check rxpck/s, txpck/s, and rxdrop/txdrop for dropped packets. - Socket Pressure (
ss -s): Check TCP socket counts. Iftimewaitorestabis in tens of thousands, connection pool exhaustion is occurring.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Use the USE Method. Run 'uptime' -> 'vmstat 1 5' -> 'free -h' -> 'iostat -xz 1 5'. High 'wa' in vmstat means Disk I/O bottleneck; high 'si/so' means memory swapping; high 'r' means CPU saturation."
⚡ 60-Second Elevator Pitch Talking Points
- Step 1: 'uptime' -> compare 1/5/15m load average against 'nproc' core count.
- Step 2: 'vmstat 1 5' -> check 'r' (CPU saturation), 'b' (I/O blocked), and 'wa' (if iowait is high, bottleneck is disk, not CPU).
- Step 3: 'free -h' -> check 'available' memory; check 'si/so' in vmstat for swap thrashing.
- Step 4: 'iostat -xz 1 5' -> check '%util' and 'await' response times for storage bottlenecks.
- Step 5: 'sar -n DEV 1 5' and 'ss -s' -> check network throughput, dropped packets, and socket saturation.
- Step 6: 'dmesg -T | tail -30' for OOM kills and hardware errors.
Advertisement