⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 188 of 194 in Kubernetes
Senior Cloud Engineer (L2) Kubernetes Storage & PVC Binding L2 Cloud Screen

Q: A Kubernetes pod is stuck in the Pending state due to a storage-related issue. What possible causes would you investigate, and how would you troubleshoot the issue?

Complete troubleshooting workflow when a Kubernetes pod is stuck in Pending state due to PersistentVolumeClaim (PVC) and CSI storage issues.

#Kubernetes #Storage #PVC #PV #CSI #Pending Pod #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When a pod remains in Pending status with storage warnings, the Kubernetes scheduler cannot place the pod onto any node because persistent storage volume constraints cannot be satisfied. I troubleshoot by checking pod events, PVC status, StorageClass provisioners, and Availability Zone topology constraints."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Inspect Pod Describe Events

Run `kubectl describe pod `. Check the Events section at the bottom. Common storage failure messages: - `0/3 nodes are available: 3 node(s) had volume node affinity conflict` - `pod has unbound immediate PersistentVolumeClaims` - `waiting for a volume to be created, either by dynamic provisioning or manual binding`

kubectl describe pod <pod-name> -n production
# Look for Events:
# Warning  FailedScheduling  pod has unbound immediate PersistentVolumeClaims
2

Investigate PersistentVolumeClaim (PVC) Binding Status

Check whether the PVC is `Bound`, `Pending`, or `Lost`. Describe the PVC to see why dynamic volume provisioning failed.

kubectl get pvc -n production
kubectl describe pvc <pvc-name> -n production
# Typical error:
# StorageClass "gp3" not found OR VolumeBindingFailed: waiting for first consumer to be created
Advertisement
3

Top Storage Failure Causes

Investigate the top 4 storage failure modes: 1. **StorageClass Misconfiguration**: StorageClass does not exist or the underlying cloud CSI driver (e.g. `ebs.csi.aws.com`) is crashed. 2. **Availability Zone Affinity Conflict**: EBS/Azure Disks are zonal resources. If an EBS volume exists in `us-east-1a`, but worker nodes with available CPU/RAM are in `us-east-1b`, the pod cannot schedule. 3. **AccessMode Mismatch**: Claim requests `ReadWriteMany` (RWX) on an EBS or Azure Disk volume that only supports `ReadWriteOnce` (RWO). 4. **Volume Limits Exceeded**: Worker node has reached maximum EBS volume attachments per EC2 instance (e.g. 28 volumes per nitro instance).

Pro Tip: StorageClass Tip: Always use volumeBindingMode: WaitForFirstConsumer in StorageClasses so volumes are provisioned in the exact Availability Zone where the pod is scheduled.
4

Check CSI Driver DaemonSet Pods

Verify that the cloud storage CSI controller and node daemonset pods are healthy in `kube-system`.

kubectl get pods -n kube-system -l app.kubernetes.io/name=aws-ebs-csi-driver
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pending storage pods are caused by unbound PVCs, missing StorageClasses, multi-attach limits, or AZ affinity conflicts. Fix by checking 'kubectl describe pvc' and using 'volumeBindingMode: WaitForFirstConsumer'."
⚡ 60-Second Elevator Pitch Talking Points
  • Run kubectl describe pod and kubectl describe pvc to identify the exact storage scheduler error.
  • Verify whether the PersistentVolume is locked in a different Availability Zone than available nodes.
  • Check AccessModes: EBS volumes support ReadWriteOnce, not ReadWriteMany.
  • Verify that the cloud CSI driver pods are healthy and running in kube-system.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →