Q: How does HPA work in Kubernetes?
Deep dive into Kubernetes HPA: metrics collection pipeline via Metrics Server and Custom Metrics API, the exact mathematical autoscaling formula, stabilization windows to prevent thrashing, and event-driven scaling with KEDA.
#Kubernetes #HPA #Autoscaling #Metrics Server #Prometheus #KEDA
🎙️ Candidate Opening & Architectural Context
"HPA automatically scales the number of pod replicas in a Deployment or StatefulSet based on observed metrics like CPU, memory, or custom business metrics. It operates as a continuous reconciliation control loop inside kube-controller-manager."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Step 1: The Control Loop & Metrics Pipeline
How HPA gathers metrics every 15 seconds (default --horizontal-pod-autoscaler-sync-period):
kubelet (cAdvisor)→Metrics Server→metrics.k8s.io API→HPA Controller→Scale ReplicaSet
- kubelet's embedded cAdvisor collects container CPU and memory usage from cgroups.
- Metrics Server scrapes kubelet summary APIs and aggregates resource usage in memory.
- The HPA controller queries
metrics.k8s.io(orcustom.metrics.k8s.iovia Prometheus Adapter) for target pods. - HPA calculates the average utilization against the pod's
spec.resources.requests(not limits!).
2️⃣
Step 2: The Exact Mathematical Formula
HPA executes this exact formula on every evaluation loop:
desiredReplicas = ceil[ currentReplicas * ( currentMetricValue / desiredMetricValue ) ]- Concrete Example: Current replicas = 3. Target CPU = 50%. Current CPU utilization = 80%.
desiredReplicas = ceil[ 3 * (80 / 50) ] = ceil[ 4.8 ] = 5 replicas.- Tolerance Window: If
currentMetricValue / desiredMetricValueis within 10% (0.9 to 1.1), HPA does not scale to prevent flapping.
3️⃣
Step 3: Flapping Prevention (Behavior) & KEDA for Events
Advanced production safeguards:
- Scale-Down Stabilization Window: Default 300s (5 minutes). HPA records the highest desired replica count over the last 5 minutes and waits before scaling down to prevent thrashing during momentary traffic dips.
- HPA Behavior Spec: Define custom scaleUp/scaleDown policies (e.g. max 100% scale up every 15s, max 10% scale down every minute).
- KEDA (Kubernetes Event-driven Autoscaling): For async workloads (SQS, Kafka, RabbitMQ), scaling on CPU is too slow. KEDA scales pods from 0 to N based on queue lag before CPU even moves.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"HPA evaluates every 15s using `desiredReplicas = ceil[currentReplicas * (currentMetric / targetMetric)]`. Crucially, target CPU percentage is calculated against the pod's resource REQUESTS, not limits. Use stabilization windows to prevent flapping, and KEDA for event-driven queue scaling."
⚡ 60-Second Elevator Pitch Talking Points
- Continuous control loop running in kube-controller-manager (polls every 15s).
- Formula: desiredReplicas = ceil[ currentReplicas * (currentMetric / targetMetric) ].
- Golden Rule: CPU percentage is relative to resource REQUESTS (not limits!). If requests aren't set, HPA cannot calculate CPU utilization.
- Flapping prevention: Uses a 5-minute stabilization window for scale-down to absorb traffic dips.
- For event queues (Kafka, SQS, RabbitMQ): Use KEDA to autoscale on queue length before CPU spikes.
Advertisement