⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 34 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Hardware Reliability & Observability Hardware Reliability
🎯 Target Role / Context: Staff AI Infrastructure Engineer managing high-density NVIDIA HGX H100 compute superclusters.

Q: An HGX H100 8-GPU node enters a degraded state where `nvidia-smi` works, but any multi-GPU training job fails immediately with 'NCCL error: unhandled system error / Fabric Manager not running'. How do you troubleshoot the NVIDIA Fabric Manager service, check NVSwitch topology, and restore intra-node NVLink fabric?

Deep triage of NVIDIA HGX H100 SXM5 systems, resolving Fabric Manager service crashes, NVSwitch link degradation, and PCIe-to-NVSwitch initialization failures in Kubernetes clusters.

#NVIDIA H100 #Fabric Manager #NVSwitch #NVLink #HGX SXM5 #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"In NVIDIA HGX H100 SXM5 systems, the 8 GPUs are not directly wired point-to-point; they are interconnected via 4 physical third-generation NVSwitch chips on the baseboard, providing 900 GB/s bidirectional NVLink bandwidth per GPU. For the NVSwitch fabric to initialize, the NVIDIA Fabric Manager daemon (`nvidia-fabricmanager`) must run on the host and match the exact driver version. When Fabric Manager fails, high-bandwidth intra-node communication collapses."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Verify Fabric Manager Service Status and Driver Version Alignment

Check the systemd Fabric Manager service: `systemctl status nvidia-fabricmanager`. A primary cause of failure is driver/package desynchronization (e.g. host kernel driver updated to 550.54.14 while Fabric Manager package remains at 550.54.15). The service will fail on startup if versions do not match character-for-character.

systemctl status nvidia-fabricmanager
journalctl -u nvidia-fabricmanager --no-pager | tail -n 20
# Error: NVLink fabric manager version (550.54.14) doesn't match driver version (550.54.15)
2

Inspect NVSwitch PCI Devices and Physical Link Errors

Check if all 4 NVSwitch devices are enumerated on the PCIe bus (`lspci -d 10de: | grep -i switch`). Query NVSwitch health using `nvidia-smi nvlink --status` and `nvswitch-audit` (or `nvidia-smi -q -d NVSWITCH`). Check for link training failures, fatal CRC errors, or hardware resets on NVSwitch ports.

# Check for all 4 NVSwitch devices on HGX H100
lspci | grep -i 'PCI bridge: NVIDIA Corporation'
# Check NVLink connection state across all 8 GPUs
nvidia-smi nvlink -s -i 0
Advertisement
3

Resolve Fabric Initialization and Remap Bad Ports via Reset Sequence

If a link is retrained in degraded state, stop all running containers. Stop Fabric Manager, unload the NVIDIA unified memory and driver modules, perform a cold PCI reset of the NVSwitch devices, and restart Fabric Manager. If fatal hardware errors persist, the baseboard NVSwitch or SXM5 socket must be flagged for RMA replacement.

systemctl stop nvidia-fabricmanager
rmmod nvidia_uvm nvidia_modeset nvidia
systemctl start nvidia-fabricmanager
# Verify fabric initialization succeeded in logs
journalctl -u nvidia-fabricmanager | grep -i "Fabric Manager initialized successfully"
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"HGX H100 SXM5 clusters rely on NVIDIA Fabric Manager to configure the physical NVSwitch chips. Any version mismatch between the kernel driver and Fabric Manager causes immediate multi-GPU collective failures."
⚡ 60-Second Elevator Pitch Talking Points
  • On HGX H100 systems, GPUs connect through 4 on-board NVSwitch chips that require the NVIDIA Fabric Manager daemon to initialize.
  • If Fabric Manager crashes or desyncs from the host driver by even a minor patch version, all multi-GPU NCCL jobs fail instantly.
  • We monitor Fabric Manager service health and NVSwitch link status in Prometheus, ensuring version locks and automated driver restarts before workloads fail.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →