Q: Your cluster has 50 nodes and pod scheduling is taking 10+ seconds. What could cause this and how do you fix it?
The kube-scheduler evaluates all nodes for each pod. At 50 nodes it shouldn't be slow unless:
#Kubernetes #Scaling & Performance #L3 #Container Orchestration #K8s #etcd
🎙️ Candidate Opening & Architectural Context
""In our production Kubernetes clusters running microservices on EKS/AKS, this was a classic operational challenge. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
The kube-scheduler evaluates all nodes for each pod. At 50 nodes it shouldn't be slow unless:
- High pod churn — many pods being created/deleted rapidly, overwhelming the scheduler queue.
- Complex affinity rules — complex pod/node affinity is O(n) per scheduling cycle.
- Scheduler config
percentageOfNodesToScore— default is 100% for small clusters but can be lowered for large ones.
2️⃣
Remediation & Permanent Safeguards
Fix: Profile with scheduler metrics, simplify affinity rules, tune percentageOfNodesToScore, optimize admission webhooks. --- ## 🟠 Advanced Scenarios
- etcd latency — scheduler reads from etcd. If etcd is slow, scheduling slows.
- Webhook admission controllers — mutating or validating webhooks add latency per-pod.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: High pod churn — many pods being created/deleted rapidly, overwhelming the scheduler queue.."
⚡ 60-Second Elevator Pitch Talking Points
- High pod churn — many pods being created/deleted rapidly, overwhelming the scheduler queue.
- Complex affinity rules — complex pod/node affinity is O(n) per scheduling cycle.
- Scheduler config percentageOfNodesToScore — default is 100% for small clusters but can be lowered...
Advertisement