Q: Your rollback script fails during an active production outage. You have 5 minutes before breaching customer SLAs. Walk me through your decision-making framework and exact actions.
High-stakes SRE decision framework when an automated rollback script fails during a production outage with only 5 minutes remaining before an enterprise SLA breach.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
T+0:00: Stop Script Debugging & Assert Incident Command
Call a hard stop on debugging the broken script. Announce on the bridge: 'We are halting rollback script triage. We have 4 minutes to SLA breach. Executing traffic diversion protocol.' Ensure one person holds command while delegates execute predefined recovery actions.
T+1:00: Route Traffic Away from the Poisoned Cluster (DNS / CDN Failover)
If running multi-region or active-standby, immediately flip Route53 ARC routing controls or Cloudflare load balancer weights to send 100% of user traffic to the healthy secondary region. If single region, switch Cloudflare/ALB to return a static maintenance page or cached response rather than throwing HTTP 500/504 errors.
# CLI Route53 Application Recovery Controller Failover
aws route53-recovery-control-config update-routing-control-states \
--routing-control-states-entries "[{\"RoutingControlArn\":\"arn:aws:...:control/primary\",\"RoutingControlState\":\"Off\"},{\"RoutingControlArn\":\"arn:aws:...:control/secondary\",\"RoutingControlState\":\"On\"}]"
T+2:30: Imperative Workload Rollback at the Compute Layer
Bypass the broken CI/CD pipeline entirely and execute an imperative rollback directly on the container orchestrator. In Kubernetes, run `kubectl rollout undo deployment/
# Bypass CI/CD and rollback deployment directly via kube-apiserver
kubectl rollout undo deployment/api-gateway -n production
kubectl rollout status deployment/api-gateway -n production --timeout=90s
T+4:30: Validate Service Health & Post-Incident Reflection
Validate 200 OK responses from synthetic canaries. Update customer status page. Once the system is stable, initiate a blameless post-mortem: analyze why the deployment failed, why the rollback script failed, and why pre-flight rollback testing was not automated in CI.
- Assert incident command and stop engineers from wasting time debugging the failed script.
- Shift traffic immediately via Route53 ARC, CDN, or load balancer to secondary clusters or warm standby.
- Bypass CI/CD pipelines to trigger imperative rollbacks directly at the cluster layer (kubectl rollout undo).
- Enable edge maintenance mode or graceful degradation to preserve SLA uptime metrics.