Q: Pod is in CrashLoopBackOff — how would you troubleshoot?
Exhaustive triage process for debugging CrashLoopBackOff: decoding container exit codes (137 OOMKill, 1 app error, 143 SIGTERM), fetching previous container logs, and fixing probe failures.
#Kubernetes #CrashLoopBackOff #OOMKilled #kubectl logs #Probes #Exit Codes
🎙️ Candidate Opening & Architectural Context
"CrashLoopBackOff means the container started, crashed, and Kubernetes is backing off before restarting it. My troubleshooting sequence always follows: Exit Code analysis -> Previous container logs -> Liveness probe evaluation."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Inspect Exit Code via kubectl describe
Run kubectl describe pod <pod-name> and examine the Last State: Terminated block:
- Exit Code 137 (OOMKilled):
Reason: OOMKilled. The Linux kernel OOM Killer terminated the container for exceeding itslimits.memory. Fix: bump memory limit or fix memory leak. - Exit Code 1 or 255 (App Error): Application code crashed (uncaught exception, syntax error, missing environment variable, failed DB connection).
- Exit Code 127 (Command Not Found): Container CMD/ENTRYPOINT executable does not exist inside the image.
- Exit Code 143 (SIGTERM): Graceful shutdown signal was received, often from a failing liveness probe or preStop hook.
- Exit Code 0 (Completed): Container finished its task and exited normally, but pod restartPolicy is
Alwaysinstead ofOnFailure/Never.
2️⃣
Fetch Previous Container Logs (--previous)
Because the crashing container has already terminated, standard logs may be empty or only show startup lines:
kubectl logs <pod-name> --previous: The golden command — reads the stdout/stderr from the crashed instance before it restarted.- If multi-container pod:
kubectl logs <pod-name> -c <container-name> --previous. - Look at the last 20 lines for stack traces, database connection timeouts, or missing config keys.
3️⃣
Check Liveness & Startup Probes
A misconfigured liveness probe will actively kill a healthy container during slow startup:
- Check events for:
Liveness probe failed: HTTP probe failed with statuscode: 500. - If an app takes 45 seconds to initialize but
initialDelaySecondsis 10, Kubernetes kills it repeatedly! - Fix: Implement a
startupProbewith generous failureThreshold, giving the app time to start before liveness checks engage.
4️⃣
Interactive Debugging with Ephemeral Containers
If logs are silent and container crashes instantly:
- Override entrypoint in a local manifest: change command to
['sh', '-c', 'sleep 3600']to keep container alive, thenkubectl exec -itto inspect files and environment. - Use
kubectl debug -it <pod-name> --image=busybox --target=<container>to attach an ephemeral debugging container sharing the process namespace.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Look at the Exit Code first (137 = OOM, 1 = App crash). Use 'kubectl logs --previous' to capture the crash stack trace. Verify startupProbe isn't killing slow-starting applications."
⚡ 60-Second Elevator Pitch Talking Points
- Run 'kubectl describe pod <name>' -> inspect Last State: Exit Code (137 = OOMKill, 1 = code exception, 127 = binary missing).
- Run 'kubectl logs <name> --previous' to retrieve the stack trace from the crashed instance.
- If Exit Code 137: Increase spec.resources.limits.memory or profile memory leak.
- Check Probes: Verify livenessProbe isn't timing out during slow boots; add startupProbe.
- If container crashes instantly: Override command with 'sleep 3600' or attach ephemeral container via 'kubectl debug'.
Advertisement