Q: A production deployment failed halfway through, leaving different service versions running. How would you recover with minimal customer impact?
Disaster recovery runbook for handling a mid-flight deployment failure that leaves inconsistent service versions running in production, handling database backward-compatibility and zero customer impact.
🛠️ Production Runbook & Step-by-Step Resolution
Freeze Deployment Pipeline & Stop Further Rollout
Immediately pause the deployment to stop the orchestrator from provisioning more broken replicas or killing remaining healthy legacy replicas.
# Pause Kubernetes deployment rollout
kubectl rollout pause deployment/payment-api -n production
Evaluate Backward Compatibility & Schema State
Check if database migrations executed. If migrations followed the Expand/Contract (Parallel Run) pattern, old and new application versions can safely run concurrently against the database without data corruption.
Execute Fast Rollback or Traffic Pinning
If the new version is throwing 5xx errors, execute an immediate rollback to the previous revision. If using Ingress/Service Mesh, route 100% of user traffic to the stable v1 subset.
kubectl rollout undo deployment/payment-api -n production
kubectl rollout status deployment/payment-api -n production
- Immediately pause the rollout (kubectl rollout pause) to prevent further replica destruction.
- Verify database backward compatibility: ensure migrations follow the Expand/Contract pattern.
- Issue an instant rollback (kubectl rollout undo) to restore 100% of traffic to the stable version.
- Review post-mortem: enforce pre-flight smoke tests, automated rollback triggers, and Blue/Green deployment.