Q: How do you design High Availability and Autoscaling in Kubernetes?
Production engineering framework for building highly available, resiliently autoscaling Kubernetes workloads: Multi-AZ node spreading, PodDisruptionBudgets (PDB), podAntiAffinity, TopologySpreadConstraints, HPA, and Karpenter integration.
#Kubernetes #High Availability #HPA #Karpenter #TopologySpreadConstraints #PDB #SRE
🎙️ Candidate Opening & Architectural Context
"I design Kubernetes HA by systematically removing single points of failure across both the control plane and data plane: Multi-AZ node groups, multiple workload replicas, PodDisruptionBudgets (PDBs), TopologySpreadConstraints across availability zones, and pod anti-affinity. Autoscaling is designed in two complementary tiers: horizontal pod scaling via HPA based on CPU/memory and custom metrics, paired with node autoscaling via Karpenter or Cluster Autoscaler to provision compute capacity dynamically."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Multi-AZ Workload Distribution & Disruption Budgets
Ensure zero downtime during node failures, AZ outages, and cluster maintenance:
# High-availability Deployment topology constraints
spec:
replicas: 6
template:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: payment-service
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: payment-service
topologyKey: kubernetes.io/hostname
- TopologySpreadConstraints: Enforce even pod distribution across availability zones using
topologyKey: topology.kubernetes.io/zonewithmaxSkew: 1. - Pod Anti-Affinity: Ensure pods of the same deployment do not land on the same physical worker node to prevent a single node crash from killing multiple replicas.
- PodDisruptionBudgets (PDB): Define
minAvailable: 60%ormaxUnavailable: 1so node draining and upgrades never violate minimum application capacity.
2️⃣
Two-Tier Autoscaling: HPA & Node Provisioning (Karpenter)
Coordinate application pod demand with physical cluster compute provisioning:
# HPA with scale-down stabilization
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: payment-hpa
namespace: prod
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: payment-service
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300
- HorizontalPodAutoscaler (HPA): Configure target CPU utilization (typically 70%) and custom business metrics (e.g. SQS queue depth or HTTP requests/sec via Prometheus Adapter).
- Stabilization Windows: Set scale-down stabilization windows (e.g., 300s) to prevent pod flapping/thrashing during erratic traffic bursts.
- Karpenter / Cluster Autoscaler: When HPA scales up pods and nodes run out of allocatable CPU/RAM, Karpenter detects the pending pods and launches right-sized EC2 nodes in under 45 seconds.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Achieve true Kubernetes HA by combining TopologySpreadConstraints across zones, podAntiAffinity across nodes, PDBs to protect during maintenance, and dual-tier autoscaling (HPA for pods, Karpenter for worker nodes)."
⚡ 60-Second Elevator Pitch Talking Points
- Distribute pod replicas across availability zones using TopologySpreadConstraints and prevent co-location with podAntiAffinity.
- Protect production capacity during node upgrades and drain events with PodDisruptionBudgets (PDB).
- Couple HPA for workload scaling with Karpenter for sub-minute node provisioning and cost-efficient bin-packing.
Advertisement