Q: Your deployment rollout is stuck. Pods from the new version aren't coming up but old ones are still running. What's happening?
This is typical RollingUpdate behavior when new pods fail healthchecks.
#Kubernetes #Troubleshooting & Debugging #L2 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""Kubernetes is a declarative desired state system; understanding the reconciliation loop is how you diagnose this quickly. The interviewer is testing: RollingUpdate strategy and rollout debugging.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
This is typical RollingUpdate behavior when new pods fail healthchecks.
- Liveness or readiness probe failing in the new version.
- New image has a bug and is crashing.
- Resource limits hit — new pods can't schedule.
2️⃣
Remediation & Permanent Safeguards
Check: kubectl rollout status deployment/ — it will show if it's stuck. Then kubectl describe pod to see why new pods aren't ready. Common causes: To rollback immediately: kubectl rollout undo deployment/ To investigate without rolling back, describe the new failing pods and check logs.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Liveness or readiness probe failing in the new version.."
⚡ 60-Second Elevator Pitch Talking Points
- Liveness or readiness probe failing in the new version.
- New image has a bug and is crashing.
- Resource limits hit — new pods can't schedule.
Advertisement