⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Kubernetes Autoscaling & Reliability Core Autoscaling

Q: How does HPA work in Kubernetes?

Deep dive into Kubernetes HPA: metrics collection pipeline via Metrics Server and Custom Metrics API, the exact mathematical autoscaling formula, stabilization windows to prevent thrashing, and event-driven scaling with KEDA.

#Kubernetes #HPA #Autoscaling #Metrics Server #Prometheus #KEDA
🎙️ Candidate Opening & Architectural Context
"HPA automatically scales the number of pod replicas in a Deployment or StatefulSet based on observed metrics like CPU, memory, or custom business metrics. It operates as a continuous reconciliation control loop inside kube-controller-manager."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Step 1: The Control Loop & Metrics Pipeline

How HPA gathers metrics every 15 seconds (default --horizontal-pod-autoscaler-sync-period):

kubelet (cAdvisor)→Metrics Server→metrics.k8s.io API→HPA Controller→Scale ReplicaSet
  • kubelet's embedded cAdvisor collects container CPU and memory usage from cgroups.
  • Metrics Server scrapes kubelet summary APIs and aggregates resource usage in memory.
  • The HPA controller queries metrics.k8s.io (or custom.metrics.k8s.io via Prometheus Adapter) for target pods.
  • HPA calculates the average utilization against the pod's spec.resources.requests (not limits!).
2️⃣

Step 2: The Exact Mathematical Formula

HPA executes this exact formula on every evaluation loop:

  • desiredReplicas = ceil[ currentReplicas * ( currentMetricValue / desiredMetricValue ) ]
  • Concrete Example: Current replicas = 3. Target CPU = 50%. Current CPU utilization = 80%.
  • desiredReplicas = ceil[ 3 * (80 / 50) ] = ceil[ 4.8 ] = 5 replicas.
  • Tolerance Window: If currentMetricValue / desiredMetricValue is within 10% (0.9 to 1.1), HPA does not scale to prevent flapping.
3️⃣

Step 3: Flapping Prevention (Behavior) & KEDA for Events

Advanced production safeguards:

  • Scale-Down Stabilization Window: Default 300s (5 minutes). HPA records the highest desired replica count over the last 5 minutes and waits before scaling down to prevent thrashing during momentary traffic dips.
  • HPA Behavior Spec: Define custom scaleUp/scaleDown policies (e.g. max 100% scale up every 15s, max 10% scale down every minute).
  • KEDA (Kubernetes Event-driven Autoscaling): For async workloads (SQS, Kafka, RabbitMQ), scaling on CPU is too slow. KEDA scales pods from 0 to N based on queue lag before CPU even moves.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"HPA evaluates every 15s using `desiredReplicas = ceil[currentReplicas * (currentMetric / targetMetric)]`. Crucially, target CPU percentage is calculated against the pod's resource REQUESTS, not limits. Use stabilization windows to prevent flapping, and KEDA for event-driven queue scaling."
⚡ 60-Second Elevator Pitch Talking Points
  • Continuous control loop running in kube-controller-manager (polls every 15s).
  • Formula: desiredReplicas = ceil[ currentReplicas * (currentMetric / targetMetric) ].
  • Golden Rule: CPU percentage is relative to resource REQUESTS (not limits!). If requests aren't set, HPA cannot calculate CPU utilization.
  • Flapping prevention: Uses a 5-minute stabilization window for scale-down to absorb traffic dips.
  • For event queues (Kafka, SQS, RabbitMQ): Use KEDA to autoscale on queue length before CPU spikes.
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes