Q: How do you implement Google's official SRE multi-window multi-burn-rate alerting model in GCP Cloud Monitoring using MQL? How do you alert on a 14.4x burn rate (2% error budget consumed in 1 hour) without waking engineers for false alarms?
Engineering production Service Level Objectives (SLOs) and Google Cloud Monitoring Query Language (MQL) multi-window multi-burn-rate alerting to eliminate alert fatigue.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Formulate Availability SLIs & Error Budgets
Define mathematical service level indicators on Google Cloud Load Balancing:
- SLI Formula:
Good Requests (HTTP < 500) / Total Valid Requests. - Target SLO: 99.9% availability over a 30-day rolling window.
- Total Error Budget: 0.1% of total requests over 30 days (100% budget = 0.001 failure rate).
Calculate Short and Long Window Burn Rates
Configure the dual-window burn rate matrix to eliminate false alarms:
- Page Alert (Critical): 14.4x burn rate consumes 2% of budget in 1 hour. Requires a 1-hour long window AND a 5-minute short window (to verify the outage is still active).
- Ticket Alert (Low Priority): 6x burn rate consumes 5% of budget in 6 hours (6-hour long window AND 30-minute short window). Sends a Jira ticket during business hours.
Write the Cloud Monitoring MQL Alert Condition
Construct the dual-window MQL query inside Cloud Monitoring:
- Fetch Metric: Fetch
loadbalancing.googleapis.com/https/request_countgrouped by response code. - MQL Ratio: Compute ratio of bad requests (5xx) over total requests across 1h and 5m windows.
- Condition:
condition (ratio_1h > 0.0144) && (ratio_5m > 0.0144).
Configure Notification Channels & Automated Runbook Links
Route alerts with contextual debugging payloads:
- Notification Channels: Routed critical burn rates to PagerDuty with auto-escalation; routed ticket burn rates to Slack and Jira.
- Documentation URL: Embedded direct links to Grafana trace explorers and GKE incident runbooks in the Cloud Monitoring alert documentation template.
- Defined 99.9% availability SLO based on Cloud Load Balancer response codes over a 30-day window.
- Implemented Google SRE dual-window alerting: 14.4x burn rate over 1h and 5m windows for PagerDuty pages.
- Constructed MQL query in Cloud Monitoring calculating error ratios simultaneously across short and long windows.
- Eliminated false alarm wake-ups while guaranteeing detection of catastrophic failures within 2 minutes.