The Production Kubernetes Troubleshooting Handbook
Over 70% of Kubernetes interview questions test your ability to debug failing workloads. This handbook provides the exact mental model, diagnostic matrix, and CLI command sequences used by Lead SREs to triage production clusters.
The Master Kubernetes Diagnostic Matrix
| Pod Status | Root Cause | Diagnostic Command | Resolution Runbook |
|---|---|---|---|
| CrashLoopBackOff | App exit error, missing secrets, failing liveness probe | kubectl logs <pod> --previous | Fix env vars, tune liveness probe initialDelaySeconds |
| OOMKilled (Exit 137) | Container exceeded cgroups memory limit | kubectl describe pod <pod> | Increase memory limits or tune JVM heap (-XX:MaxRAMPercentage) |
| Pending | Insufficient CPU/memory, node taints, PV binding failure | kubectl get events --sort-by=.metadata.creationTimestamp | Scale node groups, check Karpenter/ClusterAutoscaler |
| ImagePullBackOff | Image tag mismatch, ECR IAM auth failure, Docker Hub rate limits | kubectl describe pod <pod> | grep Events -A 10 | Verify imagePullSecrets or AWS IRSA permissions |
Essential Golden Commands for the Interview
# 1. View previous container crash logs
kubectl logs <pod-name> -c <container> --previous --tail=100
# 2. Inspect real-time scheduler and kubelet events
kubectl get events -n <namespace> --sort-by='.lastTimestamp'
# 3. Fast diagnostic exec with network tools
kubectl debug pod/<pod-name> -it --image=nicolaka/netshoot -- /bin/bash
# 4. Check cluster node resource commitments
kubectl top nodes && kubectl top pods -A --sort-by='memory'