Q: During an EKS deployment across multiple namespaces, one pod gets stuck in Pending state. Walk me through your troubleshooting approach.
Systematic diagnostic workflow when a Kubernetes deployment across multiple namespaces leaves exactly one pod stuck in Pending state.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Execute kubectl describe pod for Scheduler Verdict
Inspect the Events table in `kubectl describe pod
kubectl describe pod <pod-name> -n <namespace>
# Common output:
# 0/12 nodes are available: 3 Insufficient cpu, 5 node(s) had untolerated taint, 4 node(s) didn't match PodTopologySpread.
Check Node Affinity, Selectors & Taints/Tolerations
Check whether the failing pod spec has specific constraints that other namespaces don't: - `nodeSelector` or `nodeAffinity` targeting an instance group that doesn't exist (e.g. `topology.kubernetes.io/zone=us-east-1a`). - Missing tolerations for dedicated node pool taints (e.g. `dedicated=gpu:NoSchedule` or `spot=true:NoSchedule`).
kubectl get nodes --show-labels
kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints
Check AWS VPC CNI Subnet IP Exhaustion
On Amazon EKS using AWS VPC CNI, every pod receives a real secondary private IP from the node's VPC subnet. If one specific Availability Zone's subnet runs out of available IP addresses (`AvailableIpAddressCount: 0`), new pods targeting that subnet are stuck in Pending.
aws ec2 describe-subnets \
--subnet-ids subnet-0123456789abcdef0 \
--query "Subnets[*].[SubnetId,AvailableIpAddressCount,CidrBlock]"
Check Resource Requests vs. Cluster Autoscaler / Karpenter
Check if the pod requests excessive CPU or Memory (`requests.memory: 32Gi`) that exceeds the capacity of available node sizes. Check whether Karpenter or Cluster Autoscaler is failing to spin up new nodes due to AWS EC2 service quotas.
- Run kubectl describe pod to inspect the scheduler's rejection reasons.
- Verify nodeSelectors, nodeAffinity rules, and taints/tolerations on worker node groups.
- Check AWS VPC subnet IP availability to ensure the CNI has free IP addresses for new pods.
- Inspect Karpenter/Cluster Autoscaler logs to verify if autoscaling is blocked by AWS capacity quotas.