Q: What actually happens under the hood when a Kubernetes Pod restarts?
Deep dive into what actually happens when a Kubernetes Pod 'restarts': differentiating in-place container restarts by Kubelet from Pod deletion/rescheduling by controllers, restartPolicy exponential backoffs, and graceful shutdown lifecycles.
#Kubernetes #Pod Lifecycle #Kubelet #CrashLoopBackOff #SIGTERM #SRE
🎙️ Candidate Opening & Architectural Context
"In Kubernetes, people often say 'the pod restarted', but technically pods don't restart — containers inside the pod restart. When a container process exits or fails a liveness probe, the local Kubelet on that node detects the exit code and applies the pod's restartPolicy. If restartPolicy is Always or OnFailure, Kubelet restarts the container in-place using exponential backoff (CrashLoopBackOff), incrementing the container's restartCount while preserving the Pod's IP, UID, and ephemeral volume data. If the Pod is deleted, evicted, or rescheduled, the entire Pod is destroyed and a brand new Pod with a new UID and IP is created by the controller."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
In-Place Container Restarts & Exponential Backoff
How Kubelet handles container crashes without destroying the parent Pod sandbox:
# Inspect container restart count and termination reason
kubectl get pod api-7b8f9c-xyz -n prod -o jsonpath='{range .status.containerStatuses[*]}{.name}{" | Restarts: "}{.restartCount}{" | ExitCode: "}{.lastState.terminated.exitCode}{" | Reason: "}{.lastState.terminated.reason}{"\n"}{end}'
# Check previous container logs before the restart
kubectl logs api-7b8f9c-xyz -n prod --previous --tail=50
- Kubelet Reconciliation: The Kubelet monitors container runtime cgroups. When PID 1 inside the container exits, Kubelet cleans up the container cgroup.
- Preserved Pod Sandbox: The network namespace (Pause container), Pod IP address, and mounted volumes (emptyDir, PVCs) remain intact across container restarts.
- CrashLoopBackOff Delay: Kubelet delays restarts with an exponential backoff (10s, 20s, 40s, 80s, 160s, capping at 300s / 5 minutes) to prevent thrashing node resources.
2️⃣
Pod Replacement, OOMKills & Graceful Termination
What happens when Kubernetes terminates or replaces a Pod entirely:
# Stream recent Kubelet events for pod lifecycle and eviction signals
kubectl get events -n prod --field-selector involvedObject.name=api-7b8f9c-xyz --sort-by=.metadata.creationTimestamp
# Describe pod to review exit status and termination history
kubectl describe pod api-7b8f9c-xyz -n prod | grep -A 10 'Last State:'
- OOMKill (Exit Code 137): When a container exceeds its memory limit, the Linux kernel OOM Killer immediately sends SIGKILL (9 + 128 = 137). Kubelet records
Reason: OOMKilledand restarts the container. - Graceful Termination Sequence: Kubelet removes pod from endpoints -> executes
preStophook -> sendsSIGTERM-> waitsterminationGracePeriodSeconds(default 30s) -> sendsSIGKILL. - Controller Replacement: If a Deployment rolling update occurs or a node fails, the ReplicaSet creates a new Pod object with a fresh UID, new IP address, and newly allocated emptyDirs.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Containers restart in-place inside the existing pod sandbox (preserving Pod IP and volumes) with Kubelet exponential backoff up to 300s. A Pod itself is only replaced when a controller (ReplicaSet/Karpenter) destroys the old Pod object and schedules a new one."
⚡ 60-Second Elevator Pitch Talking Points
- Distinguish container restarts (Kubelet restarts process in-place, keeping Pod IP and storage) from Pod replacement (controller creates new UID and IP).
- Understand CrashLoopBackOff as Kubelet exponential backoff (10s to 300s) protecting the host from thrashing loops.
- Diagnose container exit codes immediately: Exit 137 indicates kernel OOMKill, while Exit 1 or 143 reflects application errors and SIGTERM.
Advertisement