Q: A pod has been running fine for weeks and suddenly starts failing with `ImagePullBackOff`. Nothing in the pod spec changed. What could cause this?
Since nothing changed in the spec, suspect external changes:
#Kubernetes #Troubleshooting & Debugging #L3 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""Kubernetes is a declarative desired state system; understanding the reconciliation loop is how you diagnose this quickly. The interviewer is testing: Image registry auth and image availability awareness.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Since nothing changed in the spec, suspect external changes:
- Registry credentials expired — imagePullSecret token rotated or expired.
- Image was deleted from the registry — someone deleted the tag from Docker Hub or ECR.
- Registry is down or unreachable — network issue or registry outage.
2️⃣
Remediation & Permanent Safeguards
Check: kubectl describe pod — the event will say exactly which registry returned what error (401 Unauthorized, 404 Not Found, etc.).
- Rate limiting — Docker Hub has pull rate limits for unauthenticated/free accounts.
- Private registry changed auth — ECR tokens expire every 12 hours if not refreshed.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Registry credentials expired — imagePullSecret token rotated or expired.."
⚡ 60-Second Elevator Pitch Talking Points
- Registry credentials expired — imagePullSecret token rotated or expired.
- Image was deleted from the registry — someone deleted the tag from Docker Hub or ECR.
- Registry is down or unreachable — network issue or registry outage.
Advertisement