Q: How do you configure Kubernetes ResourceQuotas and LimitRanges for GPU extended resources (`requests.nvidia.com/gpu`), and how do you design admission webhooks to prevent rogue users from hoarding GPUs with indefinite sleep commands?
Enforcing multi-tenant GPU quota governance, fair-share access, and priority preemption using Kubernetes Extended Resources, LimitRanges, and custom admission controllers.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Configure Namespace GPU ResourceQuotas and LimitRanges
Define Kubernetes `ResourceQuota` objects targeting extended resources: `requests.nvidia.com/gpu` and `limits.nvidia.com/gpu`. Configure `LimitRange` to mandate that any pod requesting GPUs must define explicit requests and limits that match (as Kubernetes requires requests to equal limits for GPU devices).
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-nlp-quota
namespace: nlp-research
spec:
hard:
requests.nvidia.com/gpu: "16"
limits.nvidia.com/gpu: "16"
Implement PriorityClasses and Preemption for Batch Jobs
Establish clear PriorityClasses: `system-critical` (device plugins, ingress), `production-serving` (high priority, preemption enabled), and `batch-research` (low priority, preemptible). When production inference surges, the scheduler automatically preempts low-priority training jobs, which save state and back off gracefully.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: batch-research
value: 1000
preemptionPolicy: PreemptLowerPriority
globalDefault: false
Deploy Admission Webhook to Prevent Indefinite GPU Sleep Hoarding
Deploy a validating admission webhook (using Kyverno, Gatekeeper, or custom controller). Intercept pod creation in non-production namespaces: reject pods that request GPUs but have commands containing `sleep infinity`, or mandate an `activeDeadlineSeconds` (e.g. max 8 hours for interactive debug pods) to ensure abandoned pods terminate automatically.
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: enforce-gpu-max-duration
spec:
validationFailureAction: Enforce
rules:
- name: check-active-deadline
match:
resources:
kinds: [Pod]
namespaces: ["*-research", "*-dev"]
validate:
message: "Interactive GPU pods must specify activeDeadlineSeconds <= 28800 (8h)"
pattern:
spec:
activeDeadlineSeconds: "<=28800"
- Without quotas, developers hoard expensive GPUs with indefinite sleep scripts.
- We enforce hard ResourceQuotas on `nvidia.com/gpu` per team namespace and mandate matching requests/limits via LimitRanges.
- A Kyverno admission policy enforces an 8-hour `activeDeadlineSeconds` limit on research pods, ensuring forgotten development pods automatically terminate and free up compute.