⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Linux/SRE Interview Questions Scenario 3 of 6 in Linux/SRE
Staff SRE / Principal Architect Linux/SRE Incident Command & Disaster Recovery Sev-1 Incident Command

Q: Your rollback script fails during an active production outage. You have 5 minutes before breaching customer SLAs. Walk me through your decision-making framework and exact actions.

High-stakes SRE decision framework when an automated rollback script fails during a production outage with only 5 minutes remaining before an enterprise SLA breach.

#SRE #Incident Management #Rollback #Disaster Recovery #Failover #SLA Breach
🎙️ Candidate Opening & Architectural Context
"When an automated rollback script fails during an active outage with minutes before SLA breach, trying to debug the rollback script is an anti-pattern. As Incident Commander, my singular mandate is service restoration and blast radius mitigation, not fixing deployment automation. I immediately pivot to macroscopic traffic diversion: CDN maintenance fallbacks, DNS/ALB failover to DR/standby, or shedding non-critical load."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

T+0:00: Stop Script Debugging & Assert Incident Command

Call a hard stop on debugging the broken script. Announce on the bridge: 'We are halting rollback script triage. We have 4 minutes to SLA breach. Executing traffic diversion protocol.' Ensure one person holds command while delegates execute predefined recovery actions.

Pro Tip: Incident Command Axiom: Never debug deployment tooling during an SLA-critical outage. Pivot immediately to traffic diversion or manual failover.
2

T+1:00: Route Traffic Away from the Poisoned Cluster (DNS / CDN Failover)

If running multi-region or active-standby, immediately flip Route53 ARC routing controls or Cloudflare load balancer weights to send 100% of user traffic to the healthy secondary region. If single region, switch Cloudflare/ALB to return a static maintenance page or cached response rather than throwing HTTP 500/504 errors.

# CLI Route53 Application Recovery Controller Failover
aws route53-recovery-control-config update-routing-control-states \
  --routing-control-states-entries "[{\"RoutingControlArn\":\"arn:aws:...:control/primary\",\"RoutingControlState\":\"Off\"},{\"RoutingControlArn\":\"arn:aws:...:control/secondary\",\"RoutingControlState\":\"On\"}]"
Advertisement
3

T+2:30: Imperative Workload Rollback at the Compute Layer

Bypass the broken CI/CD pipeline entirely and execute an imperative rollback directly on the container orchestrator. In Kubernetes, run `kubectl rollout undo deployment/` or patch the Deployment image directly to the prior stable Docker digest.

# Bypass CI/CD and rollback deployment directly via kube-apiserver
kubectl rollout undo deployment/api-gateway -n production
kubectl rollout status deployment/api-gateway -n production --timeout=90s
4

T+4:30: Validate Service Health & Post-Incident Reflection

Validate 200 OK responses from synthetic canaries. Update customer status page. Once the system is stable, initiate a blameless post-mortem: analyze why the deployment failed, why the rollback script failed, and why pre-flight rollback testing was not automated in CI.

Halt Script Debugging→Divert Traffic / DNS Failover→Imperative kubectl Undo→Validate Synthetic Health→Blameless Postmortem
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Under imminent SLA breach, do not debug deployment scripts. Pivot immediately to macroscopic recovery: reroute traffic to secondary regions or execute imperative orchestrator rollbacks (kubectl rollout undo)."
⚡ 60-Second Elevator Pitch Talking Points
  • Assert incident command and stop engineers from wasting time debugging the failed script.
  • Shift traffic immediately via Route53 ARC, CDN, or load balancer to secondary clusters or warm standby.
  • Bypass CI/CD pipelines to trigger imperative rollbacks directly at the cluster layer (kubectl rollout undo).
  • Enable edge maintenance mode or graceful degradation to preserve SLA uptime metrics.
Advertisement
Want more Linux/SRE scenarios?
Explore our complete collection of scenario-based Linux/SRE interview runbooks.
Browse All Linux/SRE Questions →