⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Kubernetes Advanced Scenarios Staff SRE Scenario [L3]

Q: A cluster-autoscaler is not scaling up even though pods are Pending. What could be wrong?

1. Pod is unschedulable for a reason other than resources — e.g., node affinity requires a specific label that no node type has. CA won't...

#Kubernetes #Advanced Scenarios #L3 #Container Orchestration #K8s #EC2
🎙️ Candidate Opening & Architectural Context
""When troubleshooting Kubernetes, I always follow a structured layered model: Pod status -> Events -> Logs -> Network. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Check: CA logs — kubectl logs -n kube-system — it logs exactly why it's not scaling.

  • Pod is unschedulable for a reason other than resources — e.g., node affinity requires a specific label that no node type has. CA won't add nodes it can't schedule the pod on.
  • Max node count reached — CA has a configured max. --max-nodes-total or per-node-group limit.
  • Pod has cluster-autoscaler.kubernetes.io/safe-to-evict: false — CA may refuse to scale if eviction of existing pods is blocked.
  • Cooldown period — CA has a scale-up cooldown (default 10 min). May be waiting.
2️⃣

Remediation & Permanent Safeguards

  • Budget exhausted — cloud account has hit EC2/VM quota.
  • CA can't provision the requested instance type — spot capacity unavailable.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Pod is unschedulable for a reason other than resources — e.g., node affinity requires a specific label that no node type has. CA w."
⚡ 60-Second Elevator Pitch Talking Points
  • Pod is unschedulable for a reason other than resources — e.g., node affinity requires a specific ...
  • Max node count reached — CA has a configured max. --max-nodes-total or per-node-group limit.
  • Pod has cluster-autoscaler.kubernetes.io/safe-to-evict: false — CA may refuse to scale if evictio...
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes