⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 7 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure GPU Slicing & Multi-Tenancy GPU Slicing
🎯 Target Role / Context: Staff AI Infrastructure Engineer designing multi-tenant clusters hosting diverse workloads from lightweight embeddings to LLM fine-tuning.

Q: How does NVIDIA Multi-Instance GPU (MIG) differ fundamentally from time-slicing and MPS? How do you partition an 80GB A100 into mixed MIG slices (e.g., 3g.40gb, 1g.10gb) and expose them natively to Kubernetes pods with guaranteed fault isolation?

Partitioning NVIDIA A100/H100 GPUs into physically isolated hardware instances (MIG) for secure multi-tenancy, deterministic QoS, and Kubernetes resource scheduling via GPU Operator.

#MIG #NVIDIA A100 #NVIDIA H100 #GPU Slicing #Kubernetes #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Sharing enterprise GPUs (A100/H100) among multiple teams is crucial for cost optimization. However, standard software techniques like time-slicing lack memory isolation (a single pod OOM crashes all colocated pods), and MPS (Multi-Process Service) provides no memory protection or error containment. NVIDIA MIG provides true physical hardware partitioning across compute, memory controllers, and cache."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Compare Slicing Paradigms: Time-Slicing vs MPS vs MIG

Time-slicing interleaves kernels sequentially through the GPU with zero memory isolation—one runaway container can allocate 80GB VRAM and crash neighbors. MPS provides shared memory address space without hard QoS or fault isolation. MIG physically partitions GPU SM clusters, memory crossbars, L2 cache slices, and DMA channels. If a pod in a MIG slice crashes or triggers a CUDA error, the other slices continue executing unaffected.

# Enable MIG mode on physical GPU via nvidia-smi
sudo nvidia-smi -i 0 -mig 1
2

Define MIG Geometry Profiles via NVIDIA GPU Operator ConfigMap

Rather than manually executing nvidia-smi commands on nodes, configure dynamic MIG partitioning using the NVIDIA GPU Operator's mig-parted configuration. Define geometry templates specifying how GPUs are carved up into GPU Instances (GI) and Compute Instances (CI).

# mig-config.yaml in NVIDIA GPU Operator
version: v1
mig-configs:
  mixed-profile:
    - devices: [0]
      mig-enabled: true
      mig-devices:
        "1g.10gb": 4
        "3g.40gb": 1
Advertisement
3

Expose and Schedule MIG Extended Resources in Kubernetes

Configure the GPU Operator Device Plugin with `mig.strategy: mixed`. The plugin registers dedicated resource names with Kubelet. Pods request specific slices via native resource requests. Kubelet and the NVIDIA container runtime bind the container exclusively to the designated physical MIG device node.

apiVersion: v1
kind: Pod
metadata:
  name: bert-embeddings
spec:
  containers:
  - name: embedding-service
    image: org/embed:latest
    resources:
      limits:
        nvidia.com/mig-1g.10gb: 1
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"NVIDIA MIG provides hardware-level compute and memory isolation by partitioning SMs, memory controllers, and L2 cache into distinct instances. The NVIDIA GPU Operator automates profile partitioning and exposes granular resources (e.g., `nvidia.com/mig-1g.10gb`) to Kubernetes."
⚡ 60-Second Elevator Pitch Talking Points
  • Time-slicing lacks memory isolation; one bad pod will OOM the entire physical GPU.
  • NVIDIA MIG on A100/H100 physically carves the GPU silicon into up to seven independent hardware instances with dedicated memory controllers and SMs.
  • We manage MIG profiles declaratively via the NVIDIA GPU Operator, vending precise 10GB or 40GB slices to developer pods with guaranteed QoS and zero cross-tenant interference.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →