⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 41 of 50 in AI/ML Infrastructure & GPU
Senior AI Infrastructure Engineer AI/ML Infrastructure GPU Orchestration & Kubernetes GPU Governance
🎯 Target Role / Context: Senior AI Infrastructure Engineer managing shared multi-team Kubernetes GPU clusters.

Q: How do you configure Kubernetes ResourceQuotas and LimitRanges for GPU extended resources (`requests.nvidia.com/gpu`), and how do you design admission webhooks to prevent rogue users from hoarding GPUs with indefinite sleep commands?

Enforcing multi-tenant GPU quota governance, fair-share access, and priority preemption using Kubernetes Extended Resources, LimitRanges, and custom admission controllers.

#ResourceQuota #Extended Resources #Admission Webhook #Governance #Kubernetes #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"GPUs are the most expensive resource in modern cloud infrastructure. Without strict governance, individual developers can hoard dozens of GPUs by launching pods with `sleep infinity` or requesting entire nodes without setting time limits. Establishing an equitable multi-tenant cluster requires native ResourceQuotas, priority classes, and dynamic admission validation."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Configure Namespace GPU ResourceQuotas and LimitRanges

Define Kubernetes `ResourceQuota` objects targeting extended resources: `requests.nvidia.com/gpu` and `limits.nvidia.com/gpu`. Configure `LimitRange` to mandate that any pod requesting GPUs must define explicit requests and limits that match (as Kubernetes requires requests to equal limits for GPU devices).

apiVersion: v1
kind: ResourceQuota
metadata:
  name: team-nlp-quota
  namespace: nlp-research
spec:
  hard:
    requests.nvidia.com/gpu: "16"
    limits.nvidia.com/gpu: "16"
2

Implement PriorityClasses and Preemption for Batch Jobs

Establish clear PriorityClasses: `system-critical` (device plugins, ingress), `production-serving` (high priority, preemption enabled), and `batch-research` (low priority, preemptible). When production inference surges, the scheduler automatically preempts low-priority training jobs, which save state and back off gracefully.

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: batch-research
value: 1000
preemptionPolicy: PreemptLowerPriority
globalDefault: false
Advertisement
3

Deploy Admission Webhook to Prevent Indefinite GPU Sleep Hoarding

Deploy a validating admission webhook (using Kyverno, Gatekeeper, or custom controller). Intercept pod creation in non-production namespaces: reject pods that request GPUs but have commands containing `sleep infinity`, or mandate an `activeDeadlineSeconds` (e.g. max 8 hours for interactive debug pods) to ensure abandoned pods terminate automatically.

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: enforce-gpu-max-duration
spec:
  validationFailureAction: Enforce
  rules:
  - name: check-active-deadline
    match:
      resources:
        kinds: [Pod]
        namespaces: ["*-research", "*-dev"]
    validate:
      message: "Interactive GPU pods must specify activeDeadlineSeconds <= 28800 (8h)"
      pattern:
        spec:
          activeDeadlineSeconds: "<=28800"
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Preventing GPU hoarding requires combining Kubernetes ResourceQuotas on `nvidia.com/gpu`, strict LimitRanges, PriorityClass preemption, and admission webhooks that enforce maximum pod lifespans (`activeDeadlineSeconds`)."
⚡ 60-Second Elevator Pitch Talking Points
  • Without quotas, developers hoard expensive GPUs with indefinite sleep scripts.
  • We enforce hard ResourceQuotas on `nvidia.com/gpu` per team namespace and mandate matching requests/limits via LimitRanges.
  • A Kyverno admission policy enforces an 8-hour `activeDeadlineSeconds` limit on research pods, ensuring forgotten development pods automatically terminate and free up compute.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →