Q: Pod is stuck in Pending — what would you check?
Step-by-step diagnostic workflow for pods stuck in Pending state: decoding kube-scheduler events, capacity exhaustion, taints/tolerations, PVC binding, and autoscaling response.
#Kubernetes #Scheduler #kubectl describe #Resource Limits #PVC #Taints
🎙️ Candidate Opening & Architectural Context
"A Pod stuck in Pending means the kube-scheduler cannot find a node that meets all the pod's constraints, or volume mounting / admission controllers are blocked. My primary tool is immediately 'kubectl describe pod <pod-name>'."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Inspect Scheduler Events First
Run kubectl describe pod <pod-name> and look directly at the Events section at the bottom:
- FailedScheduling message: The scheduler explains in plain English why each node was rejected (e.g.
0/6 nodes available: 3 Insufficient cpu, 3 node(s) had untolerated taint). - FailedMount / FailedAttachVolume: Storage subsystem is blocked attempting to attach or format an EBS/NFS volume.
2️⃣
Root Cause 1: Insufficient Node Capacity (CPU/Memory)
Scheduler calculates fit based on requests, not actual usage:
- Run
kubectl describe nodes | grep -A 8 'Allocated resources'to see node allocation percentages. - If pods have huge CPU/memory requests (e.g.
cpu: 4on 4-core nodes), no single node can fit them. - Fix: Adjust application requests, or ensure Cluster Autoscaler / Karpenter is provisioning new nodes.
3️⃣
Root Cause 2: Node Selectors, Affinity & Taints
Filter constraints that eliminate eligible nodes:
spec.nodeSelector: Label typo (e.g.disk: ssdwhen nodes are labeleddisktype: ssd).nodeAffinity/podAntiAffinity: Hard anti-affinity (requiredDuringSchedulingIgnoredDuringExecution) preventing pods from running on the same node/AZ.- Taints without Tolerations: Nodes tainted with
dedicated=gpu:NoScheduleor uncordoned maintenance taints.
4️⃣
Root Cause 3: Unbound PVCs & Missing Config
Check persistent storage and required configuration:
- Check PVC state:
kubectl get pvc. IfPending, check StorageClass and CSI provisioner. - Volume Multi-AZ Trap: In AWS, an EBS volume lives in
us-east-1a. If nodes inus-east-1aare full, scheduler cannot place the pod inus-east-1b. - Check referenced ConfigMaps/Secrets: If a volume references a non-existent ConfigMap, pod cannot start.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Always run 'kubectl describe pod' and read the Scheduler Events. Differentiate resource request starvation (fit calculation) from affinity/taint rules and AZ-locked EBS volume binding."
⚡ 60-Second Elevator Pitch Talking Points
- Run 'kubectl describe pod <name>' and read the Events section: FailedScheduling tells you why.
- Check Resources: verify node allocated requests ('kubectl describe nodes') vs pod spec.resources.requests.
- Check Constraints: nodeSelector, nodeAffinity, and taints/tolerations that block placement.
- Check Storage: verify PVC status ('kubectl get pvc') - check if volume is locked to a different AWS AZ.
- Verify Cluster Autoscaler / Karpenter logs to ensure nodes are actively spinning up.
- Fix: Tune requests, correct label selectors, add tolerations, or trigger node autoscaling.
Advertisement