Q: When a GPU triggers a fatal hardware error (e.g. Xid 79 'GPU fallen off the bus' or Xid 48 'Double-bit ECC error'), pods continue scheduling to the node and fail immediately. How do you engineer an automated health controller that traps Xid events, isolates the node, and initiates cloud instance replacement?
Engineering a Kubernetes self-healing controller and DaemonSet that monitors kernel NVIDIA Xid errors, automatically cordons failed nodes, safely evicts pods, and triggers cloud instance retirement.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy an eBPF / kmsg Kernel Event Trapper DaemonSet
Deploy a lightweight Go DaemonSet that opens `/dev/kmsg` or attaches an eBPF tracepoint to `nvidia:nvrm_ioctl`. Parse incoming kernel messages in real time for NVIDIA driver Xid errors. Extract the GPU PCIe bus ID and classify the failure code into non-fatal (Xid 13, 32) vs fatal hardware errors (Xid 48, 62, 79, 92).
// Go kmsg parser snippet
if strings.Contains(line, "NVRM: Xid") {
xidCode := extractXid(line)
if isFatal(xidCode) {
remediateNode(xidCode)
}
}
Immediate Node Cordoning and Taint Injection
The instant a fatal Xid is detected, the agent communicates directly with the Kubernetes API server using its local service account: it cordons the node (`spec.unschedulable: true`) and applies a hard taint (`gpu.nvidia.com/hardware-failure:NoExecute`). The `NoExecute` taint immediately evicts all running pods (triggering Kubernetes scheduler rescheduling) while preventing new workloads from landing on the broken hardware.
kubectl cordon ip-10-0-50-22.ec2.internal
kubectl taint nodes ip-10-0-50-22.ec2.internal gpu.nvidia.com/hardware-failure=xid79:NoExecute
Trigger Cloud Instance Termination via Karpenter / AWS API
For persistent hardware failures (such as uncorrectable ECC or PCIe bus drops), restarting the pod or reloading kernel drivers will not fix damaged silicon. The remediation controller emits an event to Karpenter or invokes the AWS EC2 API (`aws ec2 terminate-instances`), terminating the bad instance and provisioning a fresh, healthy GPU node within 2 minutes.
# Controller calls AWS EC2 API to terminate bad instance
aws ec2 terminate-instances --instance-ids i-0a1b2c3d4e5f67890
- When a GPU dies, standard Kubernetes considers the node healthy, scheduling pods into an immediate CrashLoop.
- Our self-healing DaemonSet monitors kernel `/dev/kmsg` for fatal NVIDIA Xid errors in real time.
- Upon detecting an uncorrectable fault (e.g. Xid 79), it immediately applies a `NoExecute` taint to evict workloads and triggers Karpenter to terminate the bad instance and provision healthy replacement silicon.