⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Junior / Associate DevOps [L1] Observability Core Fundamentals [L1]

Q: A developer says "We should alert on average CPU being high". You say "No, that's a symptom. What's the root cause?" Explain the difference between alerting on symptoms vs root causes with a concrete example.

Symptoms are resource metrics (CPU, memory, disk). Root causes are user-facing impacts (errors, latency, requests failing).

#Observability #Observability #L1 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Alert design philosophy, SRE thinking, user impact.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Symptoms are resource metrics (CPU, memory, disk). Root causes are user-facing impacts (errors, latency, requests failing).

  • Symptom: CPU > 80%
  • Root cause: Application latency > 500ms OR error rate > 1%
  • A batch job (acceptable, expected, users don't care)
  • A memory leak in a thread (critical, users see slow requests)
  • A noisy neighbor VM (critical for your app, but your app itself is fine)
2️⃣

Remediation & Permanent Safeguards

Example: You could have high CPU from: Why symptom alerts fail: If you alert on "CPU > 80%", you'll wake on-call for the batch job and ignore the memory leak causing user errors. Alert fatigue makes engineers stop responding to alerts. Root cause alerting: Instead, alert on: These align with what users *actually experience*. If the root cause is firing, you're guaranteed there's a real problem worth waking for. If CPU is high but latency and errors are normal, sleep through it.

  • Efficient code using available resources (healthy, no issue)
  • "Error rate > 1%" (users are failing)
  • "P99 latency > 500ms for 5 min" (user experience degraded)
  • "API response 500 errors > 50/min" (app crashed or hung)
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Symptom: CPU > 80%."
⚡ 60-Second Elevator Pitch Talking Points
  • Symptom: CPU > 80%
  • Root cause: Application latency > 500ms OR error rate > 1%
  • A batch job (acceptable, expected, users don't care)
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability