⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 195 of 196 in Kubernetes
Senior DevOps / SRE [L2] Kubernetes Troubleshooting & Pod Lifecycle Core Pod Mechanics
🎯 Target Role / Context: Senior Kubernetes SRE & Platform Engineer Core Question

Q: What is the Kubernetes Pod `restartPolicy`? Explain the difference between `Always`, `OnFailure`, and `Never`, how workload controllers enforce these policies, and how Kubelet manages restart backoff delays.

Comprehensive architectural guide to Kubernetes pod restartPolicy (Always, OnFailure, Never): workload controller constraints (Deployments vs Jobs), Kubelet exponential backoff delay formulas, and troubleshooting CrashLoopBackOff.

#kubernetes pod restart policy #pod restart policy #restartPolicy Always vs OnFailure vs Never #pod restart policy k8s #kubelet restart backoff #Kubernetes Job restartPolicy #CrashLoopBackOff #Kubernetes #SRE
🎙️ Candidate Opening & Architectural Context
"Understanding Kubernetes `spec.restartPolicy` is essential for diagnosing container lifecycles, Job execution, and CrashLoopBackOff incidents. The restart policy applies to all containers within a Pod and is enforced directly by the node's Kubelet daemon."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

The Three Restart Policies: Always, OnFailure, Never

How Kubelet reacts when a container inside a Pod exits:

  • Always (Default): Kubelet automatically restarts the container whenever it terminates, regardless of whether the exit code is 0 (successful completion) or non-zero (failure). Used for long-running daemon services (web servers, APIs, workers).
  • OnFailure: Kubelet restarts the container ONLY if it exits with a non-zero exit code or fails a liveness probe. If the container exits with code 0 (clean completion), Kubelet leaves it in Completed status.
  • Never: Kubelet never restarts the container under any circumstances, whether it completes successfully (code 0) or crashes (non-zero).
2️⃣

Workload Controller Rules & Validation Constraints

Kubernetes API server strictly enforces which restart policies are valid for specific controllers:

  • Deployments, StatefulSets, DaemonSets: ONLY permit restartPolicy: Always. If you attempt to set OnFailure or Never in a Deployment spec, the Kubernetes API server rejects the manifest with a validation error.
  • Jobs & CronJobs: ONLY permit restartPolicy: OnFailure or restartPolicy: Never. Jobs are designed for batch tasks that run to completion; an Always policy would cause an infinite restart loop.
  • Standalone / Bare Pods: Can configure any of the three policies (Always, OnFailure, Never).
Advertisement
3️⃣

Kubelet Exponential Restart Backoff Algorithm

How Kubelet calculates delay to prevent overwhelming node resources during repeated failures:

  • Backoff Delay Sequence: When a container crashes, Kubelet restarts it immediately on the first failure. If it crashes again, Kubelet initiates exponential backoff: 10s → 20s → 40s → 80s → 160s → 300s.
  • Maximum Cap: The backoff delay is strictly capped at 300 seconds (5 minutes).
  • Reset Condition: If the container runs stably without crashing for 10 consecutive minutes, Kubelet completely resets its backoff timer to 0.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Deployments strictly require restartPolicy: Always for persistent service uptime, while batch Jobs mandate OnFailure or Never. When repeated crashes occur, Kubelet throttles restarts up to a 5-minute cap (CrashLoopBackOff), resetting the backoff after 10 minutes of stable runtime."
⚡ 60-Second Elevator Pitch Talking Points
  • Mastered spec.restartPolicy across Always, OnFailure, and Never.
  • Highlighted API server validation constraints: Deployments permit only Always; Jobs permit only OnFailure or Never.
  • Explained Kubelet exponential backoff delay (10s to 300s cap) and its 10-minute stability reset rule.
  • Applied restart policy mechanics to troubleshoot production CrashLoopBackOff and container lifecycle failures.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →