Q: How would you use Kubernetes resource requests/limits to control costs?
Engineering methodology for eliminating 'phantom spend' in Kubernetes through container resource right-sizing, VPA recommendation mode, and intelligent node consolidation with Karpenter.
#Kubernetes #FinOps #Right-Sizing #VPA #Karpenter #Goldilocks #Bin-Packing
🎙️ Candidate Opening & Architectural Context
"In Kubernetes, you pay for what you REQUEST, not what you use. If a developer requests 4 CPUs but the container only consumes 200m, the cloud provider bills you for 4 CPUs because the scheduler reserves the space. That gap is 'phantom spend'."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
The Phantom Spend Trap (Request vs Actual)
How over-provisioned requests trigger unneeded cloud node scale-outs:
- Kube-scheduler treats `requests` as hard commitments. If node capacity is 8 vCPU, and two pods request 4 vCPU each, the node is 100% allocated.
- Cluster Autoscaler / Karpenter is forced to provision a second EC2 node, even if the actual CPU utilization on the first node is only 5%!
- Result: Cloud bills double while clusters run at 10% actual hardware utilization.
2️⃣
Automated Right-Sizing Tools (Goldilocks & VPA)
Data-driven sizing replacing developer guesswork:
- Vertical Pod Autoscaler (VPA) in 'Off' (Recommendation) Mode: Analyzes historical container usage and outputs exact recommended requests for CPU and memory.
- Fairwinds Goldilocks: Dashboard that consumes VPA recommendations and highlights over-provisioned workloads across namespaces.
- Set requests to the p95 peak utilization + 15-20% headroom, allowing pods to burst safely without blocking node scheduling.
3️⃣
Dynamic Node Consolidation with Karpenter
Maximizing node packing density:
- Enable Karpenter
consolidationPolicy: WhenUnderutilized. - Karpenter constantly evaluates cluster bin-packing: if 3 nodes are 30% full, Karpenter cordons and drains one node, moves its pods onto the remaining nodes, and terminates the empty instance in real-time.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Cloud bills scale with Kubernetes CPU/memory REQUESTS, not actual utilization. Over-provisioned requests force autoscalers to launch expensive unnecessary nodes. Right-size requests to p95 usage using VPA recommendation mode and enable Karpenter node consolidation."
⚡ 60-Second Elevator Pitch Talking Points
- Problem: Cloud billing follows requests, not usage. Over-requested pods force the cluster to spin up empty, expensive nodes.
- Solution 1: Deploy VPA in recommendation mode + Goldilocks to discover real p95 CPU/memory usage.
- Solution 2: Adjust requests down to actual p95 load + 20% buffer, freeing up node capacity.
- Solution 3: Implement Karpenter with consolidation enabled—it automatically merges sparse nodes and terminates unneeded EC2 instances.
- Result: 30–50% reduction in cluster compute spend with zero impact on application performance.
Advertisement