Q: An alerting rule fires every 3 seconds, then clears every 5 seconds, creating 150+ PagerDuty incidents per hour. The actual metric oscillates around the threshold. How do you stabilize this?
This is alert flapping—when a metric oscillates around the threshold, causing rapid alert cycles. Engineers ignore the notifications (ale...
#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Alert flapping, dampening strategies, alert fatigue reduction.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
This is alert flapping—when a metric oscillates around the threshold, causing rapid alert cycles. Engineers ignore the notifications (alert fatigue), defeating their purpose.
- Raise the evaluation window: Instead of
cpu > 80%, useavg(cpu) over 5m > 80%. Oscillations within minutes won't trigger; only sustained issues will. - Hysteresis (two-threshold approach):
- Alert fires when metric > 85% (high threshold)
- Alert clears only when metric < 75% (low threshold)
- This creates a "dead zone" between 75-85%, preventing flapping.
2️⃣
Remediation & Permanent Safeguards
Solutions (in order of increasing sophistication): Best practice example (Prometheus):
- alert: HighCPU
expr: rate(node_cpu[1m]) > 0.8
for: 5m # Must be high for 5 consecutive minutes
annotations:
summary: "CPU sustained above 80%"
- Prometheus: Use
for: 5m(must exceed threshold for 5 minutes before triggering). - Aggregation: Instead of single-instance CPU, alert on
avg(cpu) across all instances > 80%. Aggregate metrics are smoother. - Dynamic thresholding: Replace static 80% with
predict_linear(cpu[1h], 3600) > 90%. Only alert if the metric will hit 90% within the hour (giving lead time instead of flapping).
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Raise the evaluation window: Instead of cpu > 80%, use avg(cpu) over 5m > 80%. Oscillations within minutes won't trigger; only sus."
⚡ 60-Second Elevator Pitch Talking Points
- Raise the evaluation window: Instead of cpu > 80%, use avg(cpu) over 5m > 80%. Oscillations withi...
- Hysteresis (two-threshold approach):
- Alert fires when metric > 85% (high threshold)
Advertisement