Q: A container keeps restarting, but there are no application errors in the logs. What would you investigate next?
Troubleshooting checklist when a container repeatedly terminates and restarts while application logs remain completely empty or stop abruptly.
#Docker #Kubernetes #CrashLoopBackOff #OOMKilled #Barclays #Linux
🎙️ Candidate Opening & Architectural Context
"When a container restarts with empty application logs, the termination signal came from outside the application runtime—typically the Linux kernel OOM killer, Kubernetes probe failures, entrypoint shell script crashes, or cgroup memory limit violations."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1
Inspect Container Termination Exit Code & Reason
Check the exact exit code and termination reason using docker inspect or kubectl describe. Exit code 137 indicates SIGKILL (typically OOMKilled), while exit code 143 indicates SIGTERM (orchestrator shutdown).
kubectl describe pod <pod-name> | grep -E 'Last State|Exit Code|Reason'
# In Docker standalone:
docker inspect <container-id> --format 'ExitCode: {{.State.ExitCode}}, OOMKilled: {{.State.OOMKilled}}'
2
Audit Host Kernel Ring Buffer (dmesg)
Run dmesg on the host node. If the kernel Out-Of-Memory killer invoked SIGKILL, it leaves an indelible trace in dmesg showing the sacrificed process and its memory score.
dmesg -T | grep -i -E 'killed process|out of memory|oom-killer'
3
Verify Liveness Probe Configuration
Check if an aggressive liveness probe is timing out during heavy startup initialization. If initialDelaySeconds is too low, the kubelet kills the container before it finishes warming its cache.
Pro Tip: Always use startupProbe for slow-initializing applications to prevent premature livenessProbe kills.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Empty logs mean the application was killed externally. Check ExitCode 137 (OOMKilled) in docker inspect/kubectl describe and verify dmesg for kernel memory sacrifices."
⚡ 60-Second Elevator Pitch Talking Points
- Check termination Exit Code: 137 = OOMKilled (SIGKILL), 143 = orchestrator shutdown (SIGTERM), 1 = entrypoint failure.
- Run dmesg -T | grep oom-killer on the worker node to verify kernel memory exhaustion.
- Review livenessProbe thresholds: ensure startupProbe protects slow-booting applications.
- Inspect container entrypoint script for set -e exits before logging initialized.
Advertisement