Q: A Kubernetes pod is stuck in the Pending state due to a storage-related issue. What possible causes would you investigate, and how would you troubleshoot the issue?
Complete troubleshooting workflow when a Kubernetes pod is stuck in Pending state due to PersistentVolumeClaim (PVC) and CSI storage issues.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Inspect Pod Describe Events
Run `kubectl describe pod
kubectl describe pod <pod-name> -n production
# Look for Events:
# Warning FailedScheduling pod has unbound immediate PersistentVolumeClaims
Investigate PersistentVolumeClaim (PVC) Binding Status
Check whether the PVC is `Bound`, `Pending`, or `Lost`. Describe the PVC to see why dynamic volume provisioning failed.
kubectl get pvc -n production
kubectl describe pvc <pvc-name> -n production
# Typical error:
# StorageClass "gp3" not found OR VolumeBindingFailed: waiting for first consumer to be created
Top Storage Failure Causes
Investigate the top 4 storage failure modes: 1. **StorageClass Misconfiguration**: StorageClass does not exist or the underlying cloud CSI driver (e.g. `ebs.csi.aws.com`) is crashed. 2. **Availability Zone Affinity Conflict**: EBS/Azure Disks are zonal resources. If an EBS volume exists in `us-east-1a`, but worker nodes with available CPU/RAM are in `us-east-1b`, the pod cannot schedule. 3. **AccessMode Mismatch**: Claim requests `ReadWriteMany` (RWX) on an EBS or Azure Disk volume that only supports `ReadWriteOnce` (RWO). 4. **Volume Limits Exceeded**: Worker node has reached maximum EBS volume attachments per EC2 instance (e.g. 28 volumes per nitro instance).
Check CSI Driver DaemonSet Pods
Verify that the cloud storage CSI controller and node daemonset pods are healthy in `kube-system`.
kubectl get pods -n kube-system -l app.kubernetes.io/name=aws-ebs-csi-driver
- Run kubectl describe pod and kubectl describe pvc to identify the exact storage scheduler error.
- Verify whether the PersistentVolume is locked in a different Availability Zone than available nodes.
- Check AccessModes: EBS volumes support ReadWriteOnce, not ReadWriteMany.
- Verify that the cloud CSI driver pods are healthy and running in kube-system.