Q: In a dual-socket server with 8 GPUs and 8 Mellanox ConnectX-7 NICs, cross-socket UPI traffic can cut GPUDirect RDMA throughput in half. How do you configure Kubernetes Topology Manager and Device Plugins to guarantee strict PCIe tree and NUMA alignment?
Configuring Kubernetes Topology Manager, Node Feature Discovery (NFD), and CPU Manager to guarantee strict hardware locality between GPUs, CPU sockets, and Mellanox ConnectX RDMA NICs.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Inspect Hardware PCIe Tree and NUMA Locality via lstopo and nvidia-smi
Run `lstopo` and `nvidia-smi topo -m` to map the system's hardware graph. Identify the PCIe switch affinities: verify which GPU indices share PCIe root complexes (PIX / PXB) with which Mellanox network interface names (e.g. GPU 0-3 with `mlx5_0` through `mlx5_3` on NUMA 0).
# Check GPU topology matrix
nvidia-smi topo -m
# Look for NVL (NVLink), PIX (PCIe Switch), or SYS (Cross-socket UPI traversal)
Configure Kubelet Topology Manager with single-numa-node Policy
Configure Kubelet with `--cpu-manager-policy=static` and `--topology-manager-policy=single-numa-node`. Set `--topology-manager-scope=container`. When a container requests CPU, GPU, and RDMA devices, the Topology Manager queries CPU Manager and Device Manager to find a resource allocation that fits entirely within a single NUMA node. If no single NUMA node has sufficient resources, admission is rejected rather than compromised.
# /etc/kubernetes/kubelet-config.yaml
cpuManagerPolicy: static
topologyManagerPolicy: single-numa-node
topologyManagerScope: container
cpuManagerPolicyOptions:
full-pcpus-only: "true"
Expose NUMA-Aware RDMA NICs via Network Resources Injector
Deploy the Mellanox / NVIDIA Network Operator with the SR-IOV or RDMA device plugin. Configure the device plugin to advertise PCI locality alongside GPUs. Containers request both `nvidia.com/gpu: 1` and `nvidia.com/roce_nic: 1`, and Kubelet binds the container to the co-located GPU and NIC on the identical PCIe switch, enabling line-rate GPUDirect RDMA.
resources:
limits:
cpu: "16"
memory: 64Gi
nvidia.com/gpu: "4"
nvidia.com/rdma_shared_device: "4"
- Crossing CPU sockets over UPI during all-reduce collectives cuts GPUDirect RDMA throughput by up to 50%.
- We map the server's PCIe topology so each GPU is matched to its local Mellanox ConnectX NIC.
- By setting Kubelet's Topology Manager policy to `single-numa-node`, Kubernetes guarantees strict hardware locality between CPUs, GPUs, and NICs, unlocking peak 400Gbps network throughput.