⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE FinOps & Cost Container Efficiency Resource Optimization

Q: How would you use Kubernetes resource requests/limits to control costs?

Engineering methodology for eliminating 'phantom spend' in Kubernetes through container resource right-sizing, VPA recommendation mode, and intelligent node consolidation with Karpenter.

#Kubernetes #FinOps #Right-Sizing #VPA #Karpenter #Goldilocks #Bin-Packing
🎙️ Candidate Opening & Architectural Context
"In Kubernetes, you pay for what you REQUEST, not what you use. If a developer requests 4 CPUs but the container only consumes 200m, the cloud provider bills you for 4 CPUs because the scheduler reserves the space. That gap is 'phantom spend'."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

The Phantom Spend Trap (Request vs Actual)

How over-provisioned requests trigger unneeded cloud node scale-outs:

  • Kube-scheduler treats `requests` as hard commitments. If node capacity is 8 vCPU, and two pods request 4 vCPU each, the node is 100% allocated.
  • Cluster Autoscaler / Karpenter is forced to provision a second EC2 node, even if the actual CPU utilization on the first node is only 5%!
  • Result: Cloud bills double while clusters run at 10% actual hardware utilization.
2️⃣

Automated Right-Sizing Tools (Goldilocks & VPA)

Data-driven sizing replacing developer guesswork:

  • Vertical Pod Autoscaler (VPA) in 'Off' (Recommendation) Mode: Analyzes historical container usage and outputs exact recommended requests for CPU and memory.
  • Fairwinds Goldilocks: Dashboard that consumes VPA recommendations and highlights over-provisioned workloads across namespaces.
  • Set requests to the p95 peak utilization + 15-20% headroom, allowing pods to burst safely without blocking node scheduling.
3️⃣

Dynamic Node Consolidation with Karpenter

Maximizing node packing density:

  • Enable Karpenter consolidationPolicy: WhenUnderutilized.
  • Karpenter constantly evaluates cluster bin-packing: if 3 nodes are 30% full, Karpenter cordons and drains one node, moves its pods onto the remaining nodes, and terminates the empty instance in real-time.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Cloud bills scale with Kubernetes CPU/memory REQUESTS, not actual utilization. Over-provisioned requests force autoscalers to launch expensive unnecessary nodes. Right-size requests to p95 usage using VPA recommendation mode and enable Karpenter node consolidation."
⚡ 60-Second Elevator Pitch Talking Points
  • Problem: Cloud billing follows requests, not usage. Over-requested pods force the cluster to spin up empty, expensive nodes.
  • Solution 1: Deploy VPA in recommendation mode + Goldilocks to discover real p95 CPU/memory usage.
  • Solution 2: Adjust requests down to actual p95 load + 20% buffer, freeing up node capacity.
  • Solution 3: Implement Karpenter with consolidation enabled—it automatically merges sparse nodes and terminates unneeded EC2 instances.
  • Result: 30–50% reduction in cluster compute spend with zero impact on application performance.
Advertisement
Want more FinOps & Cost scenarios?
Explore our complete collection of scenario-based FinOps & Cost interview runbooks.
Browse All FinOps & Cost Questions →

📚 Related Production Scenarios in FinOps & Cost