Q: You have an alerting rule: `CPU Usage > 80% for 5 minutes`. It keeps waking you up at 3 AM for a backend batch processing worker, but it resolves itself after 10 minutes without issues. What do you do?
Waking up for non-actionable, self-resolving alerts creates alert fatigue and burns out engineers. The alert is poorly designed for this ...
#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Alert fatigue reduction, actionable alerting, understanding of batch workloads.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Waking up for non-actionable, self-resolving alerts creates alert fatigue and burns out engineers. The alert is poorly designed for this workload.
- Disable or silence the CPU alert for this specific batch worker tier.
- Replace it with a Symptom-Based Alert (SLI/SLO): Alert if the batch job queue age exceeds X minutes or if the job failure rate spikes. Alert on the *outcome* the business cares about, not the resource utilization.
2️⃣
Remediation & Permanent Safeguards
Batch processing workers are *supposed* to use 100% CPU to finish their jobs as quickly as possible. CPU usage is not a symptom of failure here; it's a measure of efficiency. I would:
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Disable or silence the CPU alert for this specific batch worker tier.."
⚡ 60-Second Elevator Pitch Talking Points
- Disable or silence the CPU alert for this specific batch worker tier.
- Replace it with a Symptom-Based Alert (SLI/SLO): Alert if the batch job queue age exceeds X minut...
Advertisement