Q: What is the Kubernetes Pod `restartPolicy`? Explain the difference between `Always`, `OnFailure`, and `Never`, how workload controllers enforce these policies, and how Kubelet manages restart backoff delays.
Comprehensive architectural guide to Kubernetes pod restartPolicy (Always, OnFailure, Never): workload controller constraints (Deployments vs Jobs), Kubelet exponential backoff delay formulas, and troubleshooting CrashLoopBackOff.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
The Three Restart Policies: Always, OnFailure, Never
How Kubelet reacts when a container inside a Pod exits:
- Always (Default): Kubelet automatically restarts the container whenever it terminates, regardless of whether the exit code is 0 (successful completion) or non-zero (failure). Used for long-running daemon services (web servers, APIs, workers).
- OnFailure: Kubelet restarts the container ONLY if it exits with a non-zero exit code or fails a liveness probe. If the container exits with code 0 (clean completion), Kubelet leaves it in Completed status.
- Never: Kubelet never restarts the container under any circumstances, whether it completes successfully (code 0) or crashes (non-zero).
Workload Controller Rules & Validation Constraints
Kubernetes API server strictly enforces which restart policies are valid for specific controllers:
- Deployments, StatefulSets, DaemonSets: ONLY permit
restartPolicy: Always. If you attempt to setOnFailureorNeverin a Deployment spec, the Kubernetes API server rejects the manifest with a validation error. - Jobs & CronJobs: ONLY permit
restartPolicy: OnFailureorrestartPolicy: Never. Jobs are designed for batch tasks that run to completion; anAlwayspolicy would cause an infinite restart loop. - Standalone / Bare Pods: Can configure any of the three policies (
Always,OnFailure,Never).
Kubelet Exponential Restart Backoff Algorithm
How Kubelet calculates delay to prevent overwhelming node resources during repeated failures:
- Backoff Delay Sequence: When a container crashes, Kubelet restarts it immediately on the first failure. If it crashes again, Kubelet initiates exponential backoff:
10s → 20s → 40s → 80s → 160s → 300s. - Maximum Cap: The backoff delay is strictly capped at 300 seconds (5 minutes).
- Reset Condition: If the container runs stably without crashing for 10 consecutive minutes, Kubelet completely resets its backoff timer to 0.
- Mastered spec.restartPolicy across Always, OnFailure, and Never.
- Highlighted API server validation constraints: Deployments permit only Always; Jobs permit only OnFailure or Never.
- Explained Kubelet exponential backoff delay (10s to 300s cap) and its 10-minute stability reset rule.
- Applied restart policy mechanics to troubleshoot production CrashLoopBackOff and container lifecycle failures.