Q: How would you debug a pod stuck in 'ContainerCreating' during deployment? What are your next 3 CLI commands?
Immediate root cause diagnosis and the exact 3 CLI commands to execute when a Kubernetes pod is stuck indefinitely in the ContainerCreating state.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Command 1: Inspect Pod Events (kubectl describe pod)
The first command reveals the exact subsystem blocking initialization: Volume Mount, Secret, ConfigMap, or Network Sandbox.
# Command 1:
kubectl describe pod <pod-name> -n <namespace>
# Look at Events table:
# MountVolume.SetUp failed for volume "pvc-storage" : rpc error ...
# FailedCreatePodSandBox: failed to setup network for sandbox ... IPAM error
Command 2: Inspect Kubelet Host Logs (journalctl -u kubelet)
If pod events are silent or generic, log into the specific node and query kubelet system logs to see the underlying CRI or CNI failure.
# Command 2:
journalctl -u kubelet --no-pager -n 50
# Reveals containerd socket errors, disk space full, or device mapper timeouts
Command 3: Inspect Containerd Runtime & CNI Plugin Logs
If kubelet is waiting on the runtime sandbox, query containerd directly via `crictl` to inspect failing container tasks and CNI allocations.
# Command 3:
crictl pods --name <pod-name>
# And check containerd logs:
journalctl -u containerd --no-pager -n 50
Top Root Causes & Fast Resolution
1. **Volume Attachment Timeout**: The previous pod replica is still holding an exclusive lock on an EBS/PersistentVolume. Manually terminate the dead pod or detach the stale cloud volume. 2. **Missing Secret/ConfigMap**: Pod spec references an optional=false Secret that does not exist in the namespace. 3. **CNI Subnet Exhaustion**: Node cannot allocate an IP address from the VPC/Calico subnet pool.
- Execute kubectl describe pod to inspect the Events log for volume mount or CNI errors.
- Run journalctl -u kubelet on the host node to see CRI runtime and sandbox errors.
- Run crictl pods and journalctl -u containerd to check container runtime timeouts directly.
- Triage common root causes: stale volume locks from previous pods, missing secrets, or CNI IP exhaustion.