Q: AWS Spot GPU instances offer up to 70% cost savings for AI training, but can be reclaimed with only 2 minutes notice. How do you design an automated spot interruption handler that traps metadata events, pauses training, takes an atomic checkpoint, and recovers the cluster?
Architecting resilient distributed training pipelines on AWS Spot GPU instances, intercepting 2-minute EC2 Spot Interruption notices, and triggering sub-minute emergency checkpointing.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Intercept 2-Minute EC2 Spot Interruption Notices
Deploy AWS Node Termination Handler (NTH) or a lightweight DaemonSet polling the EC2 Instance Metadata Service (IMDSv2) at `http://169.254.169.254/latest/meta-data/spot/instance-action` every 5 seconds. Alternatively, subscribe to Amazon EventBridge Spot Interruption Warning events via SQS.
# Poll IMDSv2 for spot interruption action
TOKEN=$(curl -s -X PUT "http://169.254.169.254/latest/api/token" -H "X-aws-ec2-metadata-token-ttl-seconds: 60")
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/spot/instance-action
# Returns: {"action": "terminate", "time": "2026-10-07T10:55:00Z"}
Signal Training Processes to Trigger Emergency Checkpoint
Upon detecting an interruption notice, the handler sends a POSIX signal (`SIGTERM` or `SIGUSR1`) to the local worker process or posts an event to an etcd/Redis coordination bus. In PyTorch, register a signal handler that catches SIGUSR1, halts the current training step, and triggers an immediate non-blocking sharded checkpoint write via TorchSnapshot directly to S3.
import signal
def emergency_checkpoint_handler(signum, frame):
print("Spot interruption notice received! Saving emergency checkpoint...")
save_sharded_checkpoint(model, optimizer, step, path="s3://checkpoints/emergency")
sys.exit(0)
signal.signal(signal.SIGUSR1, emergency_checkpoint_handler)
Automate Cluster Resumption with Karpenter and Elastic Torchrun
Karpenter immediately detects the terminating node and provisions a replacement GPU instance from a diverse pool of spot instance types and AZs. The training job uses PyTorch Elastic (Torchrun with `c10d` rendezvous backend). Once the replacement node joins, Torchrun automatically performs rendezvous, discovers the new rank cohort, and reloads the latest emergency checkpoint from S3 within 90 seconds.
# Torchrun elastic launch with automatic restart
torchrun \
--nnodes=4:8 \
--nproc_per_node=8 \
--rdzv_backend=c10d \
--rdzv_endpoint=etcd-server.internal:2379 \
--max_restarts=10 \
train.py --resume_from_checkpoint
- Spot GPUs cut training costs by 70%, but losing a node abruptly will crash standard training jobs.
- We poll the IMDSv2 metadata endpoint for the 2-minute termination notice and emit a `SIGUSR1` signal to PyTorch.
- PyTorch triggers an immediate 30-second emergency checkpoint to S3, while Karpenter provisions a replacement node and Torchrun resumes training seamlessly without human intervention.