Q: Your company's SLO for an API is 99.95% availability. The CTO asks, "Why not just set it to 99.99%?" What is the SRE argument against over-committing to higher SLOs?
Higher SLOs are exponentially more expensive to achieve, and over-committing creates perverse incentives.
#Linux #group: engineering #L3 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""During an on-call shift, our alerts triggered when a critical Linux production server exhibited this behavior. The interviewer is testing: SLO design philosophy, cost-reliability tradeoff, diminishing returns.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Higher SLOs are exponentially more expensive to achieve, and over-committing creates perverse incentives.
- Redundant everything: Multi-region active-active, automated failover, zero-downtime deployments
- Zero human-in-the-loop: 4.3 minutes is too short for humans to detect, diagnose, and fix—everything must be automated
- Dependency chain ceiling: If your database (e.g., RDS) offers 99.95% SLA, your service physically cannot achieve 99.99% unless you build multi-database redundancy
- Development velocity plummets (every change is a risk to the SLO)
2️⃣
Remediation & Permanent Safeguards
Increasing API availability from 99.9% to 99.99% reduces the allowable monthly downtime from ~43.8 minutes to just ~4.3 minutes—a 10x reduction in error budget. This leap requires fundamental architectural shifts from single-region resilience to multi-region active-active clusters with automated sub-minute failover.
- Exponential Cost Curve: Moving from three nines to four nines typically triples infrastructure costs due to multi-region data replication and networking egress.
- SRE Principle: Align SLOs with actual user pain thresholds; pursuing four nines when users are satisfied with three nines starves feature velocity.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Redundant everything: Multi-region active-active, automated failover, zero-downtime deployments."
⚡ 60-Second Elevator Pitch Talking Points
- Redundant everything: Multi-region active-active, automated failover, zero-downtime deployments
- Zero human-in-the-loop: 4.3 minutes is too short for humans to detect, diagnose, and fix—everythi...
- Dependency chain ceiling: If your database (e.g., RDS) offers 99.95% SLA, your service physically...
Advertisement