⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Observability Production Scenario [L2]

Q: An alerting rule fires every 3 seconds, then clears every 5 seconds, creating 150+ PagerDuty incidents per hour. The actual metric oscillates around the threshold. How do you stabilize this?

This is alert flapping—when a metric oscillates around the threshold, causing rapid alert cycles. Engineers ignore the notifications (ale...

#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Alert flapping, dampening strategies, alert fatigue reduction.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

This is alert flapping—when a metric oscillates around the threshold, causing rapid alert cycles. Engineers ignore the notifications (alert fatigue), defeating their purpose.

  • Raise the evaluation window: Instead of cpu > 80%, use avg(cpu) over 5m > 80%. Oscillations within minutes won't trigger; only sustained issues will.
  • Hysteresis (two-threshold approach):
  • Alert fires when metric > 85% (high threshold)
  • Alert clears only when metric < 75% (low threshold)
  • This creates a "dead zone" between 75-85%, preventing flapping.
2️⃣

Remediation & Permanent Safeguards

Solutions (in order of increasing sophistication): Best practice example (Prometheus):

- alert: HighCPU
  expr: rate(node_cpu[1m]) > 0.8
  for: 5m  # Must be high for 5 consecutive minutes
  annotations:
    summary: "CPU sustained above 80%"
  • Prometheus: Use for: 5m (must exceed threshold for 5 minutes before triggering).
  • Aggregation: Instead of single-instance CPU, alert on avg(cpu) across all instances > 80%. Aggregate metrics are smoother.
  • Dynamic thresholding: Replace static 80% with predict_linear(cpu[1h], 3600) > 90%. Only alert if the metric will hit 90% within the hour (giving lead time instead of flapping).
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Raise the evaluation window: Instead of cpu > 80%, use avg(cpu) over 5m > 80%. Oscillations within minutes won't trigger; only sus."
⚡ 60-Second Elevator Pitch Talking Points
  • Raise the evaluation window: Instead of cpu > 80%, use avg(cpu) over 5m > 80%. Oscillations withi...
  • Hysteresis (two-threshold approach):
  • Alert fires when metric > 85% (high threshold)
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability