Q: A 64-node PyTorch 70B parameter training job crashes after 3 hours with 'NCCL watchdog thread detected timeout: WorkNCCL(OpId=412, Timeout(s)=1800)'. Walk through the step-by-step diagnostic workflow to isolate whether the failure was caused by GPU hardware failure, network packet drop, or PyTorch rank desynchronization.
Root cause isolation and systematic debugging of NCCL ring initialization timeouts, watchdog hangs, and silent socket deadlocks during multi-node LLM training runs.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Enable NCCL Debug Logging and Flight Recorder
Re-run or inspect logs with NCCL telemetry flags enabled: `NCCL_DEBUG=INFO`, `NCCL_DEBUG_SUBSYS=INIT,COLL,ENV`, and `NCCL_ASYNC_ERROR_HANDLING=1`. In PyTorch 2.0+, activate the NCCL Flight Recorder (`TORCH_NCCL_RECORD_STATE_DIR=/shared/nccl_trace`), which writes an in-memory ring buffer dump of the last collective operations and identify exactly which rank failed to enter the collective.
export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=ALL
export NCCL_ASYNC_ERROR_HANDLING=1
export TORCH_NCCL_RECORD_STATE_DIR=/var/log/nccl_traces
export TORCH_NCCL_DUMP_ON_TIMEOUT=1
Analyze Host Kernel Logs and dmesg for GPU Xid Errors
Correlate the timestamp of the watchdog timeout against kernel logs (`dmesg -T`) across all 64 nodes. Look for NVIDIA Xid error codes: Xid 31 (GPU memory page fault), Xid 43 (GPU stopped processing), Xid 45 (Preemptive cleanup), or Xid 79 (GPU fallen off the bus). A single hardware Xid error immediately identifies the culprit node.
# Scan all worker nodes for recent Xid errors
ansible all_workers -m shell -a "dmesg -T | grep -E 'NVRM: Xid' | tail -n 10"
# Check for PCIe or NVLink link down
nvidia-smi nvlink --status
Isolate Fabric Packet Drops and RoCE PFC/ECN Mismatches
If all GPUs report healthy, inspect the network fabric. For RoCE v2, verify whether Priority Flow Control (PFC) storms or pause frame deadlocks occurred. Query RDMA counters (`ethtool -S eth0 | grep -E 'pfc_requests_rx|rx_pause_frames|cnp_rx'`). Run `ibnodes` and `perfquery` to detect CRC errors, symbol errors, or packet retransmissions on InfiniBand switch ports.
# Inspect RoCE RDMA counters for drops or pause frames
ethtool -S ens1f0np0 | grep -E '(rx_pfc|tx_pfc|rx_pause|tx_pause|rx_out_of_buffer)'
# Run NCCL all-reduce self-test across nodes
mpirun -np 16 --hostfile hosts /opt/nccl-tests/build/all_reduce_perf -b 8M -e 1G -f 2 -g 8
- In multi-node NCCL distributed training, every node hangs when one rank stalls.
- We enable the PyTorch NCCL Flight Recorder to immediately dump in-flight collective calls and identify the delinquent rank.
- We cross-reference kernel logs for NVIDIA Xid errors (e.g. Xid 79 or 31) and query NIC counters for RoCE PFC pause frames to separate GPU hardware lockups from network congestion.