Q: You are setting up an on-call rotation for your SRE team. What are the best practices for a healthy and sustainable on-call process?
On-call is a critical SRE function, but poorly managed on-call leads to burnout, high turnover, and slower incident response.
#Linux #group: engineering #L2 #SRE #Systems #Troubleshooting
🎙️ Candidate Opening & Architectural Context
""Never reboot a server blindly; always capture top process telemetry, lsof descriptors, and thread dumps first. The interviewer is testing: On-call practices, alert fatigue, SRE team health.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
On-call is a critical SRE function, but poorly managed on-call leads to burnout, high turnover, and slower incident response.
- Rotation structure:
- Minimum 2-person rotation (primary + secondary) to prevent single points of failure
- 1-week rotations (longer leads to burnout; shorter disrupts sleep patterns)
- Follow-the-sun model for global teams (no one is on-call during sleep hours)
- Alert quality:
- Every alert must be actionable—if the on-call can't do anything about it, it shouldn't page
- Target: fewer than 2 pages per 12-hour shift on average
- Review and eliminate noisy alerts monthly (alert fatigue kills response time)
- Use severity levels: P1 (page immediately), P2 (page during business hours), P3 (ticket only)
- Documentation and tooling:
- Every paging alert must link to a runbook
2️⃣
Remediation & Permanent Safeguards
Best practices:
- On-call engineers must have pre-provisioned access to all production systems
- Incident communication channels (Slack, PagerDuty, status page) must be set up
- Compensation and balance:
- On-call engineers receive compensation (extra pay, time off in lieu)
- After being paged overnight, the engineer gets the next day off or a late start
- Track on-call load per engineer to ensure equitable distribution
- Continuous improvement:
- Hold weekly on-call handoff meetings: outgoing on-call briefs incoming on-call
- Review every page: Was it actionable? Was the runbook sufficient? What can be automated?
- Track metrics: pages per shift, time-to-acknowledge, time-to-resolve
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Rotation structure:."
⚡ 60-Second Elevator Pitch Talking Points
- Rotation structure:
- Minimum 2-person rotation (primary + secondary) to prevent single points of failure
- 1-week rotations (longer leads to burnout; shorter disrupts sleep patterns)
Advertisement