⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Junior / Associate DevOps [L1] Kubernetes Troubleshooting & Debugging Core Fundamentals [L1]

Q: A pod is in `CrashLoopBackOff`. How do you debug it?

CrashLoopBackOff means the container starts, crashes, and Kubernetes keeps restarting it with increasing delay.

#Kubernetes #Troubleshooting & Debugging #L1 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""When troubleshooting Kubernetes, I always follow a structured layered model: Pod status -> Events -> Logs -> Network. The interviewer is testing: Log investigation and restart behavior understanding.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

CrashLoopBackOff means the container starts, crashes, and Kubernetes keeps restarting it with increasing delay.

  • kubectl logs — read the logs. If the container already restarted, use kubectl logs --previous to get logs from the last crashed instance.
  • kubectl describe pod — check exit codes. Exit code 1 = app error, 137 = OOM killed, 139 = segfault.
  • If logs are empty, the container may be crashing before writing anything — check the image and entrypoint command.
2️⃣

Remediation & Permanent Safeguards

Steps: Common causes: app error on startup, wrong config/env vars, missing secrets, OOM.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: kubectl logs — read the logs. If the container already restarted, use kubectl logs --previous to get logs from the last crashed ."
⚡ 60-Second Elevator Pitch Talking Points
  • kubectl logs — read the logs. If the container already restarted, use kubectl logs --previous t...
  • kubectl describe pod — check exit codes. Exit code 1 = app error, 137 = OOM killed, 139 = segfault.
  • If logs are empty, the container may be crashing before writing anything — check the image and en...
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes