⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE CI/CD Release Engineering Critical Incident

Q: Production deployment fails halfway — what's your rollback strategy?

Engineering strategy for recovering from halfway-failed deployments: rolling update halts, automated canary rollback, and protecting stateful databases.

#CI/CD #Rollback #Database Migrations #Canary #Blue-Green #SRE
🎙️ Candidate Opening & Architectural Context
"A deployment failing halfway means your system is in a hybrid state — some nodes or pods are running the new version while others run the old version. The strategy depends on deployment pattern and data compatibility."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Stop Forward Rollout & Contain Traffic

Halt the release immediately before more instances degrade:

  • In Kubernetes: Rolling updates halt automatically if new pods fail readiness checks or crash (governed by maxUnavailable).
  • In Canary (Argo Rollouts / Flagger): Issue an abort command (kubectl argo rollouts abort <app>) to shift 100% traffic back to the stable baseline.
  • In AWS CodeDeploy / ECS: Click Stop and roll back deployment.
2️⃣

Execute Safe Rollback

Restore all instances to the previous known-good state:

  • Run kubectl rollout undo deployment/<app> to revert to previous ReplicaSet.
  • If Blue/Green: Traffic router never switched over, or can be instantly routed back to Blue target group with 0 rebuild time.
  • Verify target health: Ensure all healthy pods are receiving traffic and serving 200 OK.
3️⃣

Verify Database & Stateful Dependencies

The most dangerous aspect of partial failures:

  • Did a migration script execute before the application containers failed?
  • If code rolled back but DB schema remains in new state: verify application is backward-compatible.
  • Golden Rule: Never deploy breaking DB schema changes and application code in the same release. Always use the Expand/Contract (Parallel Run) pattern.
4️⃣

Automated Guardrails to Prevent Recurrence

Architectural improvements for future releases:

  • Add synthetic canary health analysis: evaluate Prometheus 5xx error rate and latency over 5-minute intervals before promoting each traffic step.
  • Add pre-deployment health checks and automated database rollback scripts.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Halt forward rollout immediately. Revert traffic to the stable baseline. Never couple breaking database migrations with application deployments — use Expand/Contract."
⚡ 60-Second Elevator Pitch Talking Points
  • Halt rollout immediately: Abort canary or rolling update to prevent further degradation.
  • Revert instances: Use 'kubectl rollout undo' or flip traffic router back to 100% stable version.
  • Check Database: Verify if migrations ran; ensure schema changes follow Expand/Contract so old code remains functional.
  • Validate: Confirm healthy targets in ALB and verify application error metrics stabilize.
  • Prevent: Introduce automated Canary analysis (Argo Rollouts/Flagger) with automated rollback thresholds.
Advertisement
Want more CI/CD scenarios?
Explore our complete collection of scenario-based CI/CD interview runbooks.
Browse All CI/CD Questions →

📚 Related Production Scenarios in CI/CD