Q: A developer says "We should alert on average CPU being high". You say "No, that's a symptom. What's the root cause?" Explain the difference between alerting on symptoms vs root causes with a concrete example.
Symptoms are resource metrics (CPU, memory, disk). Root causes are user-facing impacts (errors, latency, requests failing).
#Observability #Observability #L1 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Alert design philosophy, SRE thinking, user impact.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Symptoms are resource metrics (CPU, memory, disk). Root causes are user-facing impacts (errors, latency, requests failing).
- Symptom: CPU > 80%
- Root cause: Application latency > 500ms OR error rate > 1%
- A batch job (acceptable, expected, users don't care)
- A memory leak in a thread (critical, users see slow requests)
- A noisy neighbor VM (critical for your app, but your app itself is fine)
2️⃣
Remediation & Permanent Safeguards
Example: You could have high CPU from: Why symptom alerts fail: If you alert on "CPU > 80%", you'll wake on-call for the batch job and ignore the memory leak causing user errors. Alert fatigue makes engineers stop responding to alerts. Root cause alerting: Instead, alert on: These align with what users *actually experience*. If the root cause is firing, you're guaranteed there's a real problem worth waking for. If CPU is high but latency and errors are normal, sleep through it.
- Efficient code using available resources (healthy, no issue)
- "Error rate > 1%" (users are failing)
- "P99 latency > 500ms for 5 min" (user experience degraded)
- "API response 500 errors > 50/min" (app crashed or hung)
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Symptom: CPU > 80%."
⚡ 60-Second Elevator Pitch Talking Points
- Symptom: CPU > 80%
- Root cause: Application latency > 500ms OR error rate > 1%
- A batch job (acceptable, expected, users don't care)
Advertisement