Q: HPA is scaling pods up and down too aggressively, causing instability. How do you fix it?
HPA has a stabilization window to prevent thrashing. Tune it:
#Kubernetes #Scaling & Performance #L2 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""Kubernetes is a declarative desired state system; understanding the reconciliation loop is how you diagnose this quickly. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Production Solution & Architecture
HPA has a stabilization window to prevent thrashing. Tune it: Scale up fast, scale down slow — this is the recommended pattern to handle spiky traffic without instability.
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # wait 5 min before scaling down
policies:
- type: Pods
value: 1
periodSeconds: 60 # scale down max 1 pod per minute
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Pods
value: 4
periodSeconds: 60
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: HPA has a stabilization window to prevent thrashing. Tune it:."
⚡ 60-Second Elevator Pitch Talking Points
- Immediate Triage: HPA has a stabilization window to prevent thrashing. Tune it:
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement