Q) Walk me through how you upgraded an Amazon EKS cluster from Kubernetes v1.34 to v1.36 and even after v1.37 without downtime.
Production runbook strategy for upgrading Amazon EKS clusters across minor versions (v1.34 → v1.35 → v1.36 → v1.37) with zero application downtime using sequential control plane upgrades, blue/green managed node groups, PDBs, and health validation.
Pre-checks & Compatibility Verification
Before touching the cluster, perform a complete pre-flight check of deprecated APIs and dependencies:
Upgrade One Version at a Time
Kubernetes minor version upgrades must be performed sequentially. You cannot skip minor versions:
Replace Worker Nodes (Blue/Green Node Groups)
For worker nodes, we didn't immediately terminate the existing node group:
How Did We Avoid Downtime?
Zero downtime was guaranteed because our critical microservices were engineered for High Availability:
Continuous Telemetry & Monitoring
Using CloudWatch and Prometheus/Grafana, continuously monitored throughout the upgrade:
Validate Before Removing Anything
Once all workloads were running on the new node group, we didn't immediately remove the old one:
- Conducted rigorous pre-checks: EKS Upgrade Insights, API deprecations (Pluto), add-on compatibility, and lower-environment testing.
- Upgraded strictly one minor version at a time (v1.34 → v1.35 → v1.36 → v1.37): control plane first, followed by managed add-ons.
- Implemented blue/green worker node replacement with new EKS-optimized AMI node groups; cordoned and drained nodes one-by-one.
- Guaranteed zero downtime via HA safeguards: replica counts ≥ 2, PodDisruptionBudgets, readiness probes, preStop hooks, and multi-AZ spread.
- Monitored CloudWatch/Grafana telemetry (5xx errors, latency, pending pods, ALB target health) throughout.
- Executed critical business flow smoke tests before safely decommissioning old node groups.