Q: An HGX H100 8-GPU node enters a degraded state where `nvidia-smi` works, but any multi-GPU training job fails immediately with 'NCCL error: unhandled system error / Fabric Manager not running'. How do you troubleshoot the NVIDIA Fabric Manager service, check NVSwitch topology, and restore intra-node NVLink fabric?
Deep triage of NVIDIA HGX H100 SXM5 systems, resolving Fabric Manager service crashes, NVSwitch link degradation, and PCIe-to-NVSwitch initialization failures in Kubernetes clusters.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Verify Fabric Manager Service Status and Driver Version Alignment
Check the systemd Fabric Manager service: `systemctl status nvidia-fabricmanager`. A primary cause of failure is driver/package desynchronization (e.g. host kernel driver updated to 550.54.14 while Fabric Manager package remains at 550.54.15). The service will fail on startup if versions do not match character-for-character.
systemctl status nvidia-fabricmanager
journalctl -u nvidia-fabricmanager --no-pager | tail -n 20
# Error: NVLink fabric manager version (550.54.14) doesn't match driver version (550.54.15)
Inspect NVSwitch PCI Devices and Physical Link Errors
Check if all 4 NVSwitch devices are enumerated on the PCIe bus (`lspci -d 10de: | grep -i switch`). Query NVSwitch health using `nvidia-smi nvlink --status` and `nvswitch-audit` (or `nvidia-smi -q -d NVSWITCH`). Check for link training failures, fatal CRC errors, or hardware resets on NVSwitch ports.
# Check for all 4 NVSwitch devices on HGX H100
lspci | grep -i 'PCI bridge: NVIDIA Corporation'
# Check NVLink connection state across all 8 GPUs
nvidia-smi nvlink -s -i 0
Resolve Fabric Initialization and Remap Bad Ports via Reset Sequence
If a link is retrained in degraded state, stop all running containers. Stop Fabric Manager, unload the NVIDIA unified memory and driver modules, perform a cold PCI reset of the NVSwitch devices, and restart Fabric Manager. If fatal hardware errors persist, the baseboard NVSwitch or SXM5 socket must be flagged for RMA replacement.
systemctl stop nvidia-fabricmanager
rmmod nvidia_uvm nvidia_modeset nvidia
systemctl start nvidia-fabricmanager
# Verify fabric initialization succeeded in logs
journalctl -u nvidia-fabricmanager | grep -i "Fabric Manager initialized successfully"
- On HGX H100 systems, GPUs connect through 4 on-board NVSwitch chips that require the NVIDIA Fabric Manager daemon to initialize.
- If Fabric Manager crashes or desyncs from the host driver by even a minor patch version, all multi-GPU NCCL jobs fail instantly.
- We monitor Fabric Manager service health and NVSwitch link status in Prometheus, ensuring version locks and automated driver restarts before workloads fail.