Q: How does the NVIDIA Device Plugin discover, advertise, and allocate GPUs to Kubernetes pods via NVML and Kubelet Device Manager, and how do you configure Kubelet Topology Manager to ensure strict NUMA node alignment between GPUs, CPUs, and high-speed NICs?
Architectural breakdown and troubleshooting of the NVIDIA Kubernetes Device Plugin, NVML communication, Kubelet Topology Manager, and GPU device injection into container runtimes.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Device Discovery, Health Checking, and gRPC Registration
On startup, the NVIDIA Device Plugin daemon pod initializes NVML (nvmlInit), queries all available GPUs (nvmlDeviceGetCount, nvmlDeviceGetHandleByIndex), and opens a gRPC socket at /var/lib/kubelet/device-plugins/nvidia.sock. It calls Register on Kubelet's Registration service, publishing the list of healthy GPU UUIDs as extended resources (nvidia.com/gpu). It streams continuous device health via NVML event listeners.
# Verify Device Plugin registration on node
kubectl describe node ip-10-0-45-12.ec2.internal | grep -A 8 'Allocatable:' | grep 'nvidia.com/gpu'
# Allocatable: nvidia.com/gpu: 8
Pod Scheduling and Container Device Allocation Lifecycle
When a pod requests resources (limits: nvidia.com/gpu: 4), the kube-scheduler binds the pod to a node with sufficient unallocated GPU slots. Kubelet's Device Manager executes an Allocate gRPC call to the NVIDIA Device Plugin, passing device IDs. The plugin returns container runtime CDI (Container Device Interface) specs or environment variables (NVIDIA_VISIBLE_DEVICES=GPU-UUID1,GPU-UUID2) which nvidia-container-runtime consumes to mount device nodes (/dev/nvidiaX) and NVML driver libraries into the container.
# Pod manifest requesting GPUs
spec:
containers:
- name: llama-inference
image: vllm/vllm-openai:v0.6.0
resources:
limits:
nvidia.com/gpu: '4'
memory: 64Gi
cpu: '32'
Enforce Strict NUMA Node Alignment via Kubelet Topology Manager
To prevent cross-socket UPI/QPI traffic bottlenecks during tensor parallel operations, configure Kubelet with static CPU Manager policy and Topology Manager single-numa-node policy. When allocating GPUs, Kubelet coordinates CPU Manager and Device Manager to bind worker pod CPU threads to the exact same NUMA node hosting the PCIe bus for the assigned GPUs and ConnectX InfiniBand/RoCE NICs.
# KubeletConfiguration (/etc/kubernetes/kubelet-config.yaml)
cpuManagerPolicy: static
cpuManagerPolicyOptions:
full-pcpus-only: 'true'
topologyManagerPolicy: single-numa-node
topologyManagerScope: container
- The NVIDIA Device Plugin registers extended GPU resources via gRPC with Kubelet's Device Manager using NVML.
- When pods request GPUs, nvidia-container-runtime reads CDI device specs to mount driver libraries and character devices without running privileged containers.
- We configure Kubelet's Topology Manager with `single-numa-node` policy to co-locate GPUs, CPUs, and high-speed NICs on the same NUMA socket, boosting multi-GPU throughput by up to 25%.