⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 46 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Hardware Reliability & Observability Hardware Reliability
🎯 Target Role / Context: Staff AI Infrastructure Engineer engineering automated self-healing platforms for multi-tenant GPU superclusters.

Q: When a GPU triggers a fatal hardware error (e.g. Xid 79 'GPU fallen off the bus' or Xid 48 'Double-bit ECC error'), pods continue scheduling to the node and fail immediately. How do you engineer an automated health controller that traps Xid events, isolates the node, and initiates cloud instance replacement?

Engineering a Kubernetes self-healing controller and DaemonSet that monitors kernel NVIDIA Xid errors, automatically cordons failed nodes, safely evicts pods, and triggers cloud instance retirement.

#Xid Remediation #DaemonSet #Self-Healing #Kubernetes #Node Drain #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"In a cluster with thousands of GPUs, hardware faults happen daily. When a GPU suffers a fatal failure, the node often remains 'Ready' from the perspective of standard Kubernetes Kubelet node heartbeats. New pods are scheduled to the broken node, fail instantly, and enter CrashLoopBackOff, causing cascading job failures until an engineer manually intervenes."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Deploy an eBPF / kmsg Kernel Event Trapper DaemonSet

Deploy a lightweight Go DaemonSet that opens `/dev/kmsg` or attaches an eBPF tracepoint to `nvidia:nvrm_ioctl`. Parse incoming kernel messages in real time for NVIDIA driver Xid errors. Extract the GPU PCIe bus ID and classify the failure code into non-fatal (Xid 13, 32) vs fatal hardware errors (Xid 48, 62, 79, 92).

// Go kmsg parser snippet
if strings.Contains(line, "NVRM: Xid") {
    xidCode := extractXid(line)
    if isFatal(xidCode) {
        remediateNode(xidCode)
    }
}
2

Immediate Node Cordoning and Taint Injection

The instant a fatal Xid is detected, the agent communicates directly with the Kubernetes API server using its local service account: it cordons the node (`spec.unschedulable: true`) and applies a hard taint (`gpu.nvidia.com/hardware-failure:NoExecute`). The `NoExecute` taint immediately evicts all running pods (triggering Kubernetes scheduler rescheduling) while preventing new workloads from landing on the broken hardware.

kubectl cordon ip-10-0-50-22.ec2.internal
kubectl taint nodes ip-10-0-50-22.ec2.internal gpu.nvidia.com/hardware-failure=xid79:NoExecute
Advertisement
3

Trigger Cloud Instance Termination via Karpenter / AWS API

For persistent hardware failures (such as uncorrectable ECC or PCIe bus drops), restarting the pod or reloading kernel drivers will not fix damaged silicon. The remediation controller emits an event to Karpenter or invokes the AWS EC2 API (`aws ec2 terminate-instances`), terminating the bad instance and provisioning a fresh, healthy GPU node within 2 minutes.

# Controller calls AWS EC2 API to terminate bad instance
aws ec2 terminate-instances --instance-ids i-0a1b2c3d4e5f67890
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Standard Kubelet heartbeats miss silent GPU hardware failures. An automated DaemonSet monitoring kernel kmsg for fatal Xid codes (48, 62, 79) applies immediate `NoExecute` taints and terminates degraded instances via cloud APIs."
⚡ 60-Second Elevator Pitch Talking Points
  • When a GPU dies, standard Kubernetes considers the node healthy, scheduling pods into an immediate CrashLoop.
  • Our self-healing DaemonSet monitors kernel `/dev/kmsg` for fatal NVIDIA Xid errors in real time.
  • Upon detecting an uncorrectable fault (e.g. Xid 79), it immediately applies a `NoExecute` taint to evict workloads and triggers Karpenter to terminate the bad instance and provision healthy replacement silicon.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →