⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 15 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Hardware Reliability & Observability Hardware Reliability
🎯 Target Role / Context: Staff AI Infrastructure Engineer building zero-touch self-healing infrastructure for enterprise GPU clusters.

Q: A single degraded GPU suffering uncorrectable ECC memory errors or PCIe link degradation can silently slow down or crash an entire 256-GPU distributed training run. How do you engineer a DaemonSet health agent that intercepts NVIDIA Xid errors, monitors NVML counters, and automatically cordons the node before jobs fail?

Engineering automated hardware failure detection for silent PCIe degradation, double-bit uncorrectable ECC memory errors, and kernel Xid traps with instant node cordoning and eviction.

#Xid Errors #ECC Errors #NVIDIA DCGM #Node Problem Detector #Kubernetes #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"In large GPU clusters, hardware does not always fail catastrophically with a kernel panic. Silent failures—such as PCIe links downgrading from Gen5 x16 to Gen1 x1, single-bit ECC errors escalating to uncorrectable double-bit errors (DBE), or thermal clock throttling—silently degrade collective communication or corrupt gradient updates. Detecting and cordoning degraded nodes automatically is essential for training cluster health."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Monitor Critical NVIDIA Xid Error Codes via Kernel Event Bus

Deploy a lightweight daemon or Node Problem Detector (NPD) custom plugin that tails `/dev/kmsg` for NVIDIA driver Xid errors. Categorize severity: Xid 31 (GPU memory page fault), Xid 43 (GPU execution stopped), Xid 48 (Double-bit ECC error), Xid 62 (Internal micro-controller halt), Xid 79 (GPU fallen off PCIe bus). Immediate hardware action is required for Xid 48, 62, and 79.

# Intercepting Xid events in systemd journal / kmsg
journalctl -k -g 'NVRM: Xid' -o json | jq '.'
2

Continuous NVML Hardware Health Checks and Bandwidth Probing

Run periodic (every 60s) background hardware diagnostics: query NVML for uncorrectable memory errors (`nvidia-smi -q -d ECC`), check PCIe current link generation and width (`nvidia-smi -q -d PCIE`), and monitor NVLink error counters (`nvidia-smi nvlink -e`). If PCIe link width is less than x16 or ECC uncorrectable errors exceed 0, mark the node as degraded.

# Check PCIe link status across GPUs
nvidia-smi --query-gpu=index,pci.bus_id,pcie.link.gen.current,pcie.link.gen.max,pcie.link.width.current,pcie.link.width.max --format=csv
Advertisement
3

Automated Kubernetes Cordoning and Taint Remediation

When a critical Xid or uncorrectable ECC error is detected, the daemon immediately calls the Kubernetes API: it applies `kubectl cordon ` to prevent new pods from scheduling and adds a taint (`node.kubernetes.io/gpu-unhealthy:NoSchedule`). It then notifies the distributed training operator to execute a safe checkpoint save before evicting the workload.

# Kubernetes Node Taint and Condition update via client-go
kubectl taint nodes ip-10-0-45-12.ec2.internal gpu-health=failed:NoSchedule --overwrite
kubectl cordon ip-10-0-45-12.ec2.internal
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Silent GPU degradation ruins distributed training runs. Intercepting kernel Xid errors (31, 48, 79) and tracking PCIe link bandwidth via NVML allows an automated DaemonSet to cordon unhealthy nodes before catastrophic job crashes occur."
⚡ 60-Second Elevator Pitch Talking Points
  • A single degraded GPU with PCIe link throttling or double-bit ECC errors will silently paralyze a 512-GPU training cluster.
  • We run a custom Node Problem Detector DaemonSet monitoring `/dev/kmsg` for NVIDIA Xid error codes and querying NVML for PCIe link speed downgrades.
  • Upon detecting uncorrectable errors, the agent immediately cordons the node, taints it `NoSchedule`, and triggers an automated checkpoint before initiating hardware replacement.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →