Q: What is an "Error Budget Burn Rate," and why is alerting on it superior to alerting on a static error count?
Alerting on a static threshold (e.g., "Alert if 100 errors happen") is flawed because it ignores traffic volume: 100 errors out of 100 re...
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
Alerting on a static threshold (e.g., "Alert if 100 errors happen") is flawed because it ignores traffic volume: 100 errors out of 100 requests is a furious outage; 100 errors out of 10 million requests is background noise. Burn Rate measures how fast you are consuming your 30-day Error Budget. A burn rate of 1 means you will consume exactly 100% of your budget by day 30. A burn rate of 10 implies you are consuming the budget 10 times faster than allowed and will blow the budget in 3 days. Alerting on a spike to a *Burn Rate of 10x over 1 hour* mathematically proves a severe, user-impacting outage relative to your total traffic, eliminating false positives entirely.
- Immediate Triage: Alerting on a static threshold (e.g., "Alert if 100 errors happen") is flawed because it ignore
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.