⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 37 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure GPU Orchestration & Kubernetes GPU Orchestration
🎯 Target Role / Context: Staff AI Infrastructure Engineer optimizing bare-metal GPU server architecture for extreme all-reduce efficiency.

Q: In a dual-socket server with 8 GPUs and 8 Mellanox ConnectX-7 NICs, cross-socket UPI traffic can cut GPUDirect RDMA throughput in half. How do you configure Kubernetes Topology Manager and Device Plugins to guarantee strict PCIe tree and NUMA alignment?

Configuring Kubernetes Topology Manager, Node Feature Discovery (NFD), and CPU Manager to guarantee strict hardware locality between GPUs, CPU sockets, and Mellanox ConnectX RDMA NICs.

#NUMA #PCIe #Topology Manager #ConnectX #GPUDirect RDMA #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Modern 8-GPU servers feature complex internal PCIe topologies: 4 GPUs and 4 NICs sit behind PCIe switches connected to CPU Socket 0 (NUMA Node 0), while the other 4 GPUs and 4 NICs sit behind Socket 1 (NUMA Node 1). If Kubernetes assigns GPU 0 on NUMA 0 to a pod, but pairs it with a NIC or CPU cores on NUMA 1, all tensor traffic must traverse the cross-socket Ultra Path Interconnect (UPI), causing massive latency spikes and halved bandwidth."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Inspect Hardware PCIe Tree and NUMA Locality via lstopo and nvidia-smi

Run `lstopo` and `nvidia-smi topo -m` to map the system's hardware graph. Identify the PCIe switch affinities: verify which GPU indices share PCIe root complexes (PIX / PXB) with which Mellanox network interface names (e.g. GPU 0-3 with `mlx5_0` through `mlx5_3` on NUMA 0).

# Check GPU topology matrix
nvidia-smi topo -m
# Look for NVL (NVLink), PIX (PCIe Switch), or SYS (Cross-socket UPI traversal)
2

Configure Kubelet Topology Manager with single-numa-node Policy

Configure Kubelet with `--cpu-manager-policy=static` and `--topology-manager-policy=single-numa-node`. Set `--topology-manager-scope=container`. When a container requests CPU, GPU, and RDMA devices, the Topology Manager queries CPU Manager and Device Manager to find a resource allocation that fits entirely within a single NUMA node. If no single NUMA node has sufficient resources, admission is rejected rather than compromised.

# /etc/kubernetes/kubelet-config.yaml
cpuManagerPolicy: static
topologyManagerPolicy: single-numa-node
topologyManagerScope: container
cpuManagerPolicyOptions:
  full-pcpus-only: "true"
Advertisement
3

Expose NUMA-Aware RDMA NICs via Network Resources Injector

Deploy the Mellanox / NVIDIA Network Operator with the SR-IOV or RDMA device plugin. Configure the device plugin to advertise PCI locality alongside GPUs. Containers request both `nvidia.com/gpu: 1` and `nvidia.com/roce_nic: 1`, and Kubelet binds the container to the co-located GPU and NIC on the identical PCIe switch, enabling line-rate GPUDirect RDMA.

resources:
  limits:
    cpu: "16"
    memory: 64Gi
    nvidia.com/gpu: "4"
    nvidia.com/rdma_shared_device: "4"
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Cross-socket UPI traversal cripples distributed training throughput. Configuring Kubelet's Topology Manager with `single-numa-node` policy guarantees that allocated CPUs, GPUs, and Mellanox NICs remain pinned to the identical NUMA node and PCIe switch."
⚡ 60-Second Elevator Pitch Talking Points
  • Crossing CPU sockets over UPI during all-reduce collectives cuts GPUDirect RDMA throughput by up to 50%.
  • We map the server's PCIe topology so each GPU is matched to its local Mellanox ConnectX NIC.
  • By setting Kubelet's Topology Manager policy to `single-numa-node`, Kubernetes guarantees strict hardware locality between CPUs, GPUs, and NICs, unlocking peak 400Gbps network throughput.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →