⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Kubernetes Workloads & Scheduling Core K8s Scenario

Q: Pod is stuck in Pending — what would you check?

Step-by-step diagnostic workflow for pods stuck in Pending state: decoding kube-scheduler events, capacity exhaustion, taints/tolerations, PVC binding, and autoscaling response.

#Kubernetes #Scheduler #kubectl describe #Resource Limits #PVC #Taints
🎙️ Candidate Opening & Architectural Context
"A Pod stuck in Pending means the kube-scheduler cannot find a node that meets all the pod's constraints, or volume mounting / admission controllers are blocked. My primary tool is immediately 'kubectl describe pod <pod-name>'."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Inspect Scheduler Events First

Run kubectl describe pod <pod-name> and look directly at the Events section at the bottom:

  • FailedScheduling message: The scheduler explains in plain English why each node was rejected (e.g. 0/6 nodes available: 3 Insufficient cpu, 3 node(s) had untolerated taint).
  • FailedMount / FailedAttachVolume: Storage subsystem is blocked attempting to attach or format an EBS/NFS volume.
2️⃣

Root Cause 1: Insufficient Node Capacity (CPU/Memory)

Scheduler calculates fit based on requests, not actual usage:

  • Run kubectl describe nodes | grep -A 8 'Allocated resources' to see node allocation percentages.
  • If pods have huge CPU/memory requests (e.g. cpu: 4 on 4-core nodes), no single node can fit them.
  • Fix: Adjust application requests, or ensure Cluster Autoscaler / Karpenter is provisioning new nodes.
3️⃣

Root Cause 2: Node Selectors, Affinity & Taints

Filter constraints that eliminate eligible nodes:

  • spec.nodeSelector: Label typo (e.g. disk: ssd when nodes are labeled disktype: ssd).
  • nodeAffinity / podAntiAffinity: Hard anti-affinity (requiredDuringSchedulingIgnoredDuringExecution) preventing pods from running on the same node/AZ.
  • Taints without Tolerations: Nodes tainted with dedicated=gpu:NoSchedule or uncordoned maintenance taints.
4️⃣

Root Cause 3: Unbound PVCs & Missing Config

Check persistent storage and required configuration:

  • Check PVC state: kubectl get pvc. If Pending, check StorageClass and CSI provisioner.
  • Volume Multi-AZ Trap: In AWS, an EBS volume lives in us-east-1a. If nodes in us-east-1a are full, scheduler cannot place the pod in us-east-1b.
  • Check referenced ConfigMaps/Secrets: If a volume references a non-existent ConfigMap, pod cannot start.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Always run 'kubectl describe pod' and read the Scheduler Events. Differentiate resource request starvation (fit calculation) from affinity/taint rules and AZ-locked EBS volume binding."
⚡ 60-Second Elevator Pitch Talking Points
  • Run 'kubectl describe pod <name>' and read the Events section: FailedScheduling tells you why.
  • Check Resources: verify node allocated requests ('kubectl describe nodes') vs pod spec.resources.requests.
  • Check Constraints: nodeSelector, nodeAffinity, and taints/tolerations that block placement.
  • Check Storage: verify PVC status ('kubectl get pvc') - check if volume is locked to a different AWS AZ.
  • Verify Cluster Autoscaler / Karpenter logs to ensure nodes are actively spinning up.
  • Fix: Tune requests, correct label selectors, add tolerations, or trigger node autoscaling.
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes