Q: Your AKS cluster running batch data processing, machine learning training, and test environments accounts for $40,000/month in cloud compute bills. How do you integrate Azure Spot Node Pools to reduce compute costs by 60-80% without failing mission-critical batch workflows when Azure reclaims spot capacity?
Engineering a resilient cost-optimization architecture on AKS utilizing Azure Spot Node Pools with eviction rate metrics, Kubernetes taints/tolerations, and automated node drain hooks.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Provision Dedicated AKS Spot Node Pool with Price-Capacity Eviction
Add an elastic Spot node pool to the existing production cluster:
- CLI Pool Creation: Provisioned spot pool with
az aks nodepool add --resource-group rg-aks --cluster-name aks-prod --name spotpool --priority Spot --eviction-policy Delete --spot-max-price -1 --enable-cluster-autoscaler --min-count 0 --max-count 50 --node-vm-size Standard_D8s_v5 --node-taints kubernetes.azure.com/scalesetpriority=spot:NoSchedule. - Spot Max Price: Configured
--spot-max-price -1to indicate willingness to pay up to standard on-demand price, minimizing eviction frequency to pure capacity reclaim events.
Configure Tolerations and Node Affinity on Fault-Tolerant Workloads
Explicitly schedule batch and non-critical services onto Spot infrastructure:
- Toleration Block: Added toleration in pod spec for key
kubernetes.azure.com/scalesetprioritywith operatorEqualand valuespot. - Preferred Node Affinity: Configured
preferredDuringSchedulingIgnoredDuringExecutiontargeting spot nodes, allowing pods to fall back to regular on-demand nodes if no Spot capacity is available in the region.
Deploy Azure Scheduled Events Node Termination Handler
Gracefully intercept Azure's 30-second pre-eviction notice to drain pods safely:
- Termination Handler DaemonSet: Deployed
azure-scheduled-events-handlerquerying the IMDS endpoint (http://169.254.169.254/metadata/scheduledevents). - Cordon & Drain: Upon detecting a Preempt scheduled event, the controller immediately marks the node unschedulable (
cordon) and triggers a graceful pod drain with--ignore-daemonsets --delete-emptydir-data.
Validate Financial Savings & Workload Completion SLOs
Quantify compute cost reduction and monitor job failure metrics:
- Cost Impact: Cut batch processing compute cost from $42,000/month to $11,800/month (72% net savings).
- SLO Health: Batch job completion rate remained steady at 99.98% due to automated Kubernetes Job retries.
- Add dedicated AKS Spot node pools with built-in NoSchedule taints and autoscaling.
- Target fault-tolerant workloads to Spot nodes using soft preferredDuringScheduling affinity.
- Deploy an IMDS Scheduled Events handler to intercept 30-second eviction notices and gracefully drain pods.
- Achieve over 70% compute savings on batch, CI, and test workloads with zero impact on production SLAs.