Q: How do you decide between a rollback vs a hotfix during a high-severity production incident?
Structured SRE decision framework for choosing between an immediate rollback versus a forward hotfix during high-severity production incidents: evaluating blast radius, stateful database migration risks, feature flag killswitches, and MTTR.
#CI/CD #SRE #Incident Management #Rollback #Hotfix #MTTR #Disaster Recovery
🎙️ Candidate Opening & Architectural Context
"During a Sev-1 production incident, my golden rule is: optimize for Mean Time to Recovery (MTTR) first, not for root cause analysis. I default to an immediate rollback when the incident is clearly correlated with a recent release, the change is stateless, and rollback is low-risk and well-understood. Conversely, I choose a forward hotfix or feature-flag mitigation only when an irreversible database migration has already run, when data formats have diverged, or when rolling back would cause greater customer data loss."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
The SRE Incident Decision Matrix: Rollback vs Hotfix
Four critical criteria evaluated by the Incident Commander:
# Fast rollback in Kubernetes
kubectl rollout history deployment/api -n prod
kubectl rollout undo deployment/api -n prod
kubectl rollout status deployment/api -n prod --timeout=120s
# Fast feature flag mitigation via ConfigMap or environment variable
kubectl set env deployment/api FEATURE_NEW_PAYMENTS=false -n prod
- 1. Correlation to Recent Change: Did error rates or latency spike immediately following a deployment? If yes, rollback is the primary candidate.
- 2. Data & Schema Compatibility: Has a non-backward-compatible database migration executed? If rolling back app code will cause crashes on newly formatted DB tables, rollback is dangerous — forward hotfix is required.
- 3. Feature Flag Killswitch: Can the broken feature be deactivated in seconds via LaunchDarkly or ConfigMap toggle? If yes, toggle the flag before executing any code deployment.
- 4. Estimated MTTR: Rollback takes <3 minutes via
kubectl rollout undoor Helm. A code hotfix takes 20-40 minutes (writing code, PR review, CI pipeline run). If rollback is safe, never wait on a hotfix.
2️⃣
Incident Commander Execution & Verification Protocol
Maintaining command discipline and single-threaded ownership during the outage:
# Verify error rate recovery following rollback
kubectl logs deployment/api -n prod --since=5m | grep -iE 'error|exception' | wc -l
kubectl top pods -n prod -l app=api
- Single Recovery Owner: The Incident Commander assigns one engineer to execute the rollback/hotfix while other engineers gather diagnostic logs and communicate with stakeholders.
- No Concurrent Changes: Never execute multiple emergency config changes simultaneously, as this obscures the root cause and creates conflicting failure modes.
- Metric Verification: Confirm recovery using user-facing error rate metrics in Grafana before declaring incident mitigation.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Optimize for MTTR first: default to immediate rollback for stateless changes, use feature flag killswitches for quick mitigation, and forward-fix only when irreversible database schema changes prevent clean rollbacks."
⚡ 60-Second Elevator Pitch Talking Points
- Default to immediate rollback for stateless deployments where recovery takes under 3 minutes.
- Use feature-flag killswitches to disable broken code paths instantly without redeploying containers.
- Choose a forward hotfix only when irreversible database migrations or data transformations prevent a safe rollback.
Advertisement