⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 17 of 50 in AI/ML Infrastructure & GPU
Senior AI Infrastructure Engineer AI/ML Infrastructure FinOps & Spot Resilience Spot Resilience
🎯 Target Role / Context: Senior AI Platform Engineer driving cloud FinOps and training cost optimization without sacrificing model delivery deadlines.

Q: AWS Spot GPU instances offer up to 70% cost savings for AI training, but can be reclaimed with only 2 minutes notice. How do you design an automated spot interruption handler that traps metadata events, pauses training, takes an atomic checkpoint, and recovers the cluster?

Architecting resilient distributed training pipelines on AWS Spot GPU instances, intercepting 2-minute EC2 Spot Interruption notices, and triggering sub-minute emergency checkpointing.

#Spot GPU #AWS EC2 Spot #Interruption Warning #Fault Tolerance #FinOps #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Large-scale AI model fine-tuning and pre-training represent a massive cloud expenditure. Running on Spot instances slashes infrastructure spend dramatically, but a single reclaimed instance causes a standard PyTorch DDP/FSDP job to crash with connection reset errors, losing hours of progress. Building a resilient spot platform requires proactive interruption detection and rapid checkpointing."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Intercept 2-Minute EC2 Spot Interruption Notices

Deploy AWS Node Termination Handler (NTH) or a lightweight DaemonSet polling the EC2 Instance Metadata Service (IMDSv2) at `http://169.254.169.254/latest/meta-data/spot/instance-action` every 5 seconds. Alternatively, subscribe to Amazon EventBridge Spot Interruption Warning events via SQS.

# Poll IMDSv2 for spot interruption action
TOKEN=$(curl -s -X PUT "http://169.254.169.254/latest/api/token" -H "X-aws-ec2-metadata-token-ttl-seconds: 60")
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/spot/instance-action
# Returns: {"action": "terminate", "time": "2026-10-07T10:55:00Z"}
2

Signal Training Processes to Trigger Emergency Checkpoint

Upon detecting an interruption notice, the handler sends a POSIX signal (`SIGTERM` or `SIGUSR1`) to the local worker process or posts an event to an etcd/Redis coordination bus. In PyTorch, register a signal handler that catches SIGUSR1, halts the current training step, and triggers an immediate non-blocking sharded checkpoint write via TorchSnapshot directly to S3.

import signal

def emergency_checkpoint_handler(signum, frame):
    print("Spot interruption notice received! Saving emergency checkpoint...")
    save_sharded_checkpoint(model, optimizer, step, path="s3://checkpoints/emergency")
    sys.exit(0)

signal.signal(signal.SIGUSR1, emergency_checkpoint_handler)
Advertisement
3

Automate Cluster Resumption with Karpenter and Elastic Torchrun

Karpenter immediately detects the terminating node and provisions a replacement GPU instance from a diverse pool of spot instance types and AZs. The training job uses PyTorch Elastic (Torchrun with `c10d` rendezvous backend). Once the replacement node joins, Torchrun automatically performs rendezvous, discovers the new rank cohort, and reloads the latest emergency checkpoint from S3 within 90 seconds.

# Torchrun elastic launch with automatic restart
torchrun \
  --nnodes=4:8 \
  --nproc_per_node=8 \
  --rdzv_backend=c10d \
  --rdzv_endpoint=etcd-server.internal:2379 \
  --max_restarts=10 \
  train.py --resume_from_checkpoint
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Spot GPU resilience relies on intercepting IMDSv2 interruption notices, triggering rapid sub-minute sharded checkpoints via SIGUSR1 handlers, and leveraging Torchrun elastic rendezvous to automatically resume training on replacement nodes."
⚡ 60-Second Elevator Pitch Talking Points
  • Spot GPUs cut training costs by 70%, but losing a node abruptly will crash standard training jobs.
  • We poll the IMDSv2 metadata endpoint for the 2-minute termination notice and emit a `SIGUSR1` signal to PyTorch.
  • PyTorch triggers an immediate 30-second emergency checkpoint to S3, while Karpenter provisions a replacement node and Torchrun resumes training seamlessly without human intervention.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →