⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE CI/CD Deployment Strategies & Rollbacks Production Incident

Q: How do you decide between a rollback vs a hotfix during a high-severity production incident?

Structured SRE decision framework for choosing between an immediate rollback versus a forward hotfix during high-severity production incidents: evaluating blast radius, stateful database migration risks, feature flag killswitches, and MTTR.

#CI/CD #SRE #Incident Management #Rollback #Hotfix #MTTR #Disaster Recovery
🎙️ Candidate Opening & Architectural Context
"During a Sev-1 production incident, my golden rule is: optimize for Mean Time to Recovery (MTTR) first, not for root cause analysis. I default to an immediate rollback when the incident is clearly correlated with a recent release, the change is stateless, and rollback is low-risk and well-understood. Conversely, I choose a forward hotfix or feature-flag mitigation only when an irreversible database migration has already run, when data formats have diverged, or when rolling back would cause greater customer data loss."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

The SRE Incident Decision Matrix: Rollback vs Hotfix

Four critical criteria evaluated by the Incident Commander:

# Fast rollback in Kubernetes
kubectl rollout history deployment/api -n prod
kubectl rollout undo deployment/api -n prod
kubectl rollout status deployment/api -n prod --timeout=120s

# Fast feature flag mitigation via ConfigMap or environment variable
kubectl set env deployment/api FEATURE_NEW_PAYMENTS=false -n prod
  • 1. Correlation to Recent Change: Did error rates or latency spike immediately following a deployment? If yes, rollback is the primary candidate.
  • 2. Data & Schema Compatibility: Has a non-backward-compatible database migration executed? If rolling back app code will cause crashes on newly formatted DB tables, rollback is dangerous — forward hotfix is required.
  • 3. Feature Flag Killswitch: Can the broken feature be deactivated in seconds via LaunchDarkly or ConfigMap toggle? If yes, toggle the flag before executing any code deployment.
  • 4. Estimated MTTR: Rollback takes <3 minutes via kubectl rollout undo or Helm. A code hotfix takes 20-40 minutes (writing code, PR review, CI pipeline run). If rollback is safe, never wait on a hotfix.
2️⃣

Incident Commander Execution & Verification Protocol

Maintaining command discipline and single-threaded ownership during the outage:

# Verify error rate recovery following rollback
kubectl logs deployment/api -n prod --since=5m | grep -iE 'error|exception' | wc -l
kubectl top pods -n prod -l app=api
  • Single Recovery Owner: The Incident Commander assigns one engineer to execute the rollback/hotfix while other engineers gather diagnostic logs and communicate with stakeholders.
  • No Concurrent Changes: Never execute multiple emergency config changes simultaneously, as this obscures the root cause and creates conflicting failure modes.
  • Metric Verification: Confirm recovery using user-facing error rate metrics in Grafana before declaring incident mitigation.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Optimize for MTTR first: default to immediate rollback for stateless changes, use feature flag killswitches for quick mitigation, and forward-fix only when irreversible database schema changes prevent clean rollbacks."
⚡ 60-Second Elevator Pitch Talking Points
  • Default to immediate rollback for stateless deployments where recovery takes under 3 minutes.
  • Use feature-flag killswitches to disable broken code paths instantly without redeploying containers.
  • Choose a forward hotfix only when irreversible database migrations or data transformations prevent a safe rollback.
Advertisement
Want more CI/CD scenarios?
Explore our complete collection of scenario-based CI/CD interview runbooks.
Browse All CI/CD Questions →

📚 Related Production Scenarios in CI/CD