⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 36 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure FinOps & Cost Optimization GPU FinOps
🎯 Target Role / Context: Staff AI Infrastructure Engineer leading infrastructure governance and cloud cost efficiency.

Q: A team reserves 32 H100 GPUs at $3.50/hr each ($80k/month), but their model code is bottlenecked on CPU preprocessing, resulting in only 8% actual GPU SM utilization. How do you design an automated GPU FinOps system that tracks allocated cost vs utilized cost and enforces reclamation policies?

Engineering a GPU FinOps platform combining OpenCost/Kubecost with NVIDIA DCGM hardware metrics to enforce departmental cost allocation, identify zombie GPU pods, and eliminate idle waste.

#FinOps #OpenCost #Kubecost #DCGM #GPU Allocation #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"In Kubernetes, cost allocation tools like Kubecost and OpenCost calculate costs based on requested resources (`nvidia.com/gpu: 1`). However, this hides massive financial waste: a team that requests 8 GPUs and lets them sit at 0% compute is charged the same as a team driving 95% Tensor Core utilization. True AI FinOps requires correlating financial billing with actual hardware utilization metrics."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Deploy OpenCost / Kubecost with Custom GPU Cost Profiles

Deploy OpenCost across all AI clusters. Define precise hourly amortization rates for GPU instance types (e.g. AWS p4de.24xlarge = $40.96/hr, p5.48xlarge = $98.32/hr). Tag every pod and namespace with mandatory cost center labels (`cost-center: computer-vision`, `project: llama-rag`).

# OpenCost configuration mapping GPU costs per node type
node_costs:
  "p4de.24xlarge": 40.96
  "p5.48xlarge": 98.32
  "g5.12xlarge": 5.67
2

Correlate Financial Allocation with DCGM Hardware SM Utilization

Join OpenCost financial allocation metrics with Prometheus DCGM metrics: `DCGM_FI_PROF_SM_ACTIVE`. Calculate the 'Financial Efficiency Ratio': `Efficiency = (DCGM_SM_ACTIVE / 100) * Allocated_Cost`. This reveals the exact dollar waste of underutilized GPUs per engineering department.

# PromQL calculating wasted GPU spend per namespace per month:
sum(opencost_gpu_cost_hourly * (1 - (DCGM_FI_PROF_SM_ACTIVE / 100))) by (namespace) * 730
Advertisement
3

Automate Zombie GPU Eviction and Downscaling Alerts

Deploy a governance controller that monitors running GPU pods. If a pod maintains `< 5% DCGM_FI_PROF_SM_ACTIVE` for more than 4 consecutive hours in non-production namespaces, emit an automated Slack alert to the team owner. If unacknowledged within 2 hours, automatically evict the pod and release the node back to Karpenter for termination.

# Zombie GPU alert rule
- alert: ZombieGPUAllocation
  expr: avg_over_time(DCGM_FI_PROF_SM_ACTIVE[4h]) < 5 and on(pod) kube_pod_status_phase{phase="Running"} == 1
  for: 30m
  labels: { severity: high }
  annotations:
    summary: "Pod {{ $labels.pod }} in {{ $labels.namespace }} has wasted GPU capacity (<5% SM for 4h)"
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Standard cost tools only show requested GPU costs. Correlating OpenCost billing with DCGM `SM_ACTIVE` metrics exposes true dollar waste and enables automated eviction of zombie GPU allocations."
⚡ 60-Second Elevator Pitch Talking Points
  • Teams easily spend $100k/month reserving GPUs that sit idle due to CPU bottlenecks.
  • We integrated OpenCost with NVIDIA DCGM hardware metrics to compute the exact dollar cost of underutilized GPU cycles per department.
  • Automated governance controllers flag pods running under 5% SM utilization and safely evict zombie workloads, reclaiming over $350k annually in cloud spend.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →