Q: HPA refuses to scale even though Prometheus shows CPU > 80%. Diagnose with cloud + K8s metrics.
Systematic triage guide for diagnosing why a Kubernetes Horizontal Pod Autoscaler fails to trigger replica scale-outs despite external Prometheus dashboards alerting on high CPU utilization.
#Kubernetes #HPA #Prometheus #Metrics Server #Autoscaling #SRE #Cloud
🎙️ Candidate Opening & Architectural Context
"During a flash sale, monitoring alerts page the on-call engineer: application CPU usage is sustained at 85% across all pods, customer latency is degrading, but the Deployment remains pinned at its minimum replica count of 3. The engineer reports that Prometheus shows the cluster on fire, but HPA refuses to scale out."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Understanding the Fundamental Decoupling: Prometheus vs Metrics Server
Kubernetes HPA does NOT read Prometheus by default. It queries the Kubernetes Metrics API (`metrics.k8s.io`):
- Prometheus collects metrics asynchronously via scraping agents (Node Exporter, cAdvisor).
- HPA relies on
metrics-serverscraping kubelet summary APIs every 15–60 seconds. - If
metrics-serveris degraded, failing TLS verification, or hitting API limits, HPA is blinded.
2️⃣
Execute Diagnostic CLI Triage
Inspect HPA conditions, resource definitions, and metrics API health:
# 1. Check HPA status and conditions
kubectl describe hpa <hpa-name>
# Look for:
# Conditions:
# AbleToScale: True
# ScalingActive: False (FailedGetResourceMetric)
# Current: <unknown> / 80%
# 2. Check if metrics-server is serving metrics
kubectl top pods -l app=<app-name>
kubectl get apiservice v1beta1.metrics.k8s.io
# 3. Check container resources specification
kubectl get deployment <app-name> -o yaml | grep -A 8 resources
- Missing CPU Requests: If
resources.requests.cpuis omitted, HPA calculation is mathematically impossible. HPA computes target percentage as:(actual usage / requested CPU) * 100. Without requests, HPA shows<unknown>. - Max Replicas Reached: Check if current replicas ==
spec.maxReplicas. - Stabilization Window: Check if downscale/upscale stabilization windows are throttling scaling actions.
3️⃣
Check Node Capacity & Cluster Autoscaler Blockades
If HPA updated `desiredReplicas` but pods cannot be scheduled:
kubectl get pods -l app=<app-name> | grep Pending
kubectl describe pod <pending-pod> | grep -A 5 Events
# Look for: "0/12 nodes are available: 12 Insufficient cpu."
- If nodes are fully packed and the Cloud Cluster Autoscaler / Karpenter is blocked by AWS EC2 quota limits or subnet IP exhaustion, pods remain stuck in `Pending`.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"HPA scaling failures almost always boil down to three root causes: missing container CPU requests, metrics-server failure, or cluster capacity constraints blocking Pending pods."
⚡ 60-Second Elevator Pitch Talking Points
- HPA does not read Prometheus by default; it reads the Kubernetes Metrics API (metrics-server). First, I run 'kubectl describe hpa' to check the ScalingActive condition.
- If HPA shows '<unknown> / 80%', the most common cause is that the container specification lacks 'resources.requests.cpu', making percentage calculation mathematically undefined.
- Second, I check if metrics-server is healthy using 'kubectl get apiservice v1beta1.metrics.k8s.io' and 'kubectl top pods'.
- Third, I verify whether HPA has already hit 'maxReplicas', or if upscale stabilization windows are suppressing new events.
- Finally, if desired replicas increased but actual pods are stuck in Pending, I inspect Cluster Autoscaler / Karpenter logs for EC2 quota or subnet IP exhaustion.
Advertisement