⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Kubernetes Troubleshooting & Debugging Staff SRE Scenario [L3]

Q: A node in your cluster shows `NotReady`. Your team is panicking because several services are on it. What's your action plan?

1. Don't panic — check if pods already rescheduled. Kubernetes evicts pods from NotReady nodes after pod-eviction-timeout (default 5 min)...

#Kubernetes #Troubleshooting & Debugging #L3 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""When troubleshooting Kubernetes, I always follow a structured layered model: Pod status -> Events -> Logs -> Network. The interviewer is testing: Incident response, node troubleshooting, pod eviction understanding.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

  • Don't panic — check if pods already rescheduled. Kubernetes evicts pods from NotReady nodes after pod-eviction-timeout (default 5 min). Check kubectl get pods -A -o wide | grep .
  • Cordon the node — kubectl cordon prevents new pods from scheduling there while you investigate.
  • SSH into the node and check:
  • systemctl status kubelet — is kubelet running?
  • journalctl -u kubelet -n 100 — kubelet logs.
2️⃣

Remediation & Permanent Safeguards

Execute the resolution runbook and verify workload health:

  • Disk space: df -h. Full disk is a common cause.
  • Memory: free -m.
  • Check the node's conditions: kubectl describe node — look for MemoryPressure, DiskPressure, PIDPressure.
  • If unrecoverable, drain and delete: kubectl drain --ignore-daemonsets --delete-emptydir-data then terminate the VM.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Don't panic — check if pods already rescheduled. Kubernetes evicts pods from NotReady nodes after pod-eviction-timeout (default 5 ."
⚡ 60-Second Elevator Pitch Talking Points
  • Don't panic — check if pods already rescheduled. Kubernetes evicts pods from NotReady nodes after...
  • Cordon the node — kubectl cordon prevents new pods from scheduling there while you investigate.
  • SSH into the node and check:
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes