Q: A single degraded GPU suffering uncorrectable ECC memory errors or PCIe link degradation can silently slow down or crash an entire 256-GPU distributed training run. How do you engineer a DaemonSet health agent that intercepts NVIDIA Xid errors, monitors NVML counters, and automatically cordons the node before jobs fail?
Engineering automated hardware failure detection for silent PCIe degradation, double-bit uncorrectable ECC memory errors, and kernel Xid traps with instant node cordoning and eviction.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Monitor Critical NVIDIA Xid Error Codes via Kernel Event Bus
Deploy a lightweight daemon or Node Problem Detector (NPD) custom plugin that tails `/dev/kmsg` for NVIDIA driver Xid errors. Categorize severity: Xid 31 (GPU memory page fault), Xid 43 (GPU execution stopped), Xid 48 (Double-bit ECC error), Xid 62 (Internal micro-controller halt), Xid 79 (GPU fallen off PCIe bus). Immediate hardware action is required for Xid 48, 62, and 79.
# Intercepting Xid events in systemd journal / kmsg
journalctl -k -g 'NVRM: Xid' -o json | jq '.'
Continuous NVML Hardware Health Checks and Bandwidth Probing
Run periodic (every 60s) background hardware diagnostics: query NVML for uncorrectable memory errors (`nvidia-smi -q -d ECC`), check PCIe current link generation and width (`nvidia-smi -q -d PCIE`), and monitor NVLink error counters (`nvidia-smi nvlink -e`). If PCIe link width is less than x16 or ECC uncorrectable errors exceed 0, mark the node as degraded.
# Check PCIe link status across GPUs
nvidia-smi --query-gpu=index,pci.bus_id,pcie.link.gen.current,pcie.link.gen.max,pcie.link.width.current,pcie.link.width.max --format=csv
Automated Kubernetes Cordoning and Taint Remediation
When a critical Xid or uncorrectable ECC error is detected, the daemon immediately calls the Kubernetes API: it applies `kubectl cordon
# Kubernetes Node Taint and Condition update via client-go
kubectl taint nodes ip-10-0-45-12.ec2.internal gpu-health=failed:NoSchedule --overwrite
kubectl cordon ip-10-0-45-12.ec2.internal
- A single degraded GPU with PCIe link throttling or double-bit ECC errors will silently paralyze a 512-GPU training cluster.
- We run a custom Node Problem Detector DaemonSet monitoring `/dev/kmsg` for NVIDIA Xid error codes and querying NVML for PCIe link speed downgrades.
- Upon detecting uncorrectable errors, the agent immediately cordons the node, taints it `NoSchedule`, and triggers an automated checkpoint before initiating hardware replacement.