⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Observability Staff SRE Scenario [L3]

Q: Your CPU alert threshold is 90%. The server CPU oscillates between 89% and 92% every few seconds. This causes PagerDuty to trigger the alert, resolve it, and trigger it again 50 times an hour. How do you fix this?

This is called a Flapping Alert. To fix it, you introduce Hysteresis or a Pending Duration.

#Observability #Observability #L3 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Our SRE team tackled this monitoring and metrics bottleneck to eliminate false-positive alert fatigue. The interviewer is testing: Flapping alerts, hysteresis, `for` durations in PromQL.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Production Solution & Architecture

This is called a Flapping Alert. To fix it, you introduce Hysteresis or a Pending Duration. In Prometheus, this is solved using the for clause in the alert evaluation rule. Rather than firing the millisecond the CPU hits 91%, the metric must *sustainably remain* above 90% for a continuous, unbroken 5-minute window. If it drops to 89% at minute 4, the timer resets. This guarantees you are only paged for sustained actual load, completely eliminating noise from instantaneous spikes.

alert: HighCPU
expr: cpu_usage > 90
for: 5m
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: This is called a Flapping Alert. To fix it, you introduce Hysteresis or a Pending Duration.."
⚡ 60-Second Elevator Pitch Talking Points
  • Immediate Triage: This is called a Flapping Alert. To fix it, you introduce Hysteresis or a Pending Duration.
  • Run targeted verification commands before modifying configuration.
  • Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability