Q: How does NVIDIA Multi-Instance GPU (MIG) differ fundamentally from time-slicing and MPS? How do you partition an 80GB A100 into mixed MIG slices (e.g., 3g.40gb, 1g.10gb) and expose them natively to Kubernetes pods with guaranteed fault isolation?
Partitioning NVIDIA A100/H100 GPUs into physically isolated hardware instances (MIG) for secure multi-tenancy, deterministic QoS, and Kubernetes resource scheduling via GPU Operator.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Compare Slicing Paradigms: Time-Slicing vs MPS vs MIG
Time-slicing interleaves kernels sequentially through the GPU with zero memory isolation—one runaway container can allocate 80GB VRAM and crash neighbors. MPS provides shared memory address space without hard QoS or fault isolation. MIG physically partitions GPU SM clusters, memory crossbars, L2 cache slices, and DMA channels. If a pod in a MIG slice crashes or triggers a CUDA error, the other slices continue executing unaffected.
# Enable MIG mode on physical GPU via nvidia-smi
sudo nvidia-smi -i 0 -mig 1
Define MIG Geometry Profiles via NVIDIA GPU Operator ConfigMap
Rather than manually executing nvidia-smi commands on nodes, configure dynamic MIG partitioning using the NVIDIA GPU Operator's mig-parted configuration. Define geometry templates specifying how GPUs are carved up into GPU Instances (GI) and Compute Instances (CI).
# mig-config.yaml in NVIDIA GPU Operator
version: v1
mig-configs:
mixed-profile:
- devices: [0]
mig-enabled: true
mig-devices:
"1g.10gb": 4
"3g.40gb": 1
Expose and Schedule MIG Extended Resources in Kubernetes
Configure the GPU Operator Device Plugin with `mig.strategy: mixed`. The plugin registers dedicated resource names with Kubelet. Pods request specific slices via native resource requests. Kubelet and the NVIDIA container runtime bind the container exclusively to the designated physical MIG device node.
apiVersion: v1
kind: Pod
metadata:
name: bert-embeddings
spec:
containers:
- name: embedding-service
image: org/embed:latest
resources:
limits:
nvidia.com/mig-1g.10gb: 1
- Time-slicing lacks memory isolation; one bad pod will OOM the entire physical GPU.
- NVIDIA MIG on A100/H100 physically carves the GPU silicon into up to seven independent hardware instances with dedicated memory controllers and SMs.
- We manage MIG profiles declaratively via the NVIDIA GPU Operator, vending precise 10GB or 40GB slices to developer pods with guaranteed QoS and zero cross-tenant interference.