Q: Walk me through how you upgraded an Amazon EKS cluster from Kubernetes v1.34 to v1.36 and even after v1.37 without downtime.
Production runbook strategy for upgrading Amazon EKS clusters across minor versions (v1.34 → v1.35 → v1.36 → v1.37) with zero application downtime using sequential control plane upgrades, blue/green managed node groups, PDBs, and health validation.
🛠️ Production Runbook & Step-by-Step Resolution
Pre-checks & Compatibility Verification
Before touching the cluster, perform a complete pre-flight check of deprecated APIs and dependencies:
- EKS Upgrade Insights: Checked automated AWS insights for deprecated APIs and cluster readiness.
- Deprecated Kubernetes APIs: Audited manifests and Helm charts using API deprecation tools (e.g., Pluto / kubent).
- Helm chart and application compatibility: Verified all third-party charts, CRDs, and controllers support the target version.
- Core Add-ons: Checked compatibility matrices for VPC CNI, CoreDNS, and kube-proxy for target versions.
- AWS Load Balancer Controller: Verified controller version, IAM policies (IRSA), and TargetGroupBinding CRDs.
- PDBs and replica counts: Verified PodDisruptionBudgets and replica counts (≥ 2) across all critical services.
- Monitoring components: Ensured Prometheus, Grafana, and CloudWatch were operational to capture real-time telemetry.
Upgrade One Version at a Time
Kubernetes minor version upgrades must be performed sequentially. You cannot skip minor versions:
- For each version, I first upgraded the EKS control plane (AWS handles the multi-AZ control plane upgrade without API downtime).
- Once the control plane was healthy, I validated the cluster and then upgraded the required EKS add-ons (VPC CNI, CoreDNS, kube-proxy, EBS CSI driver).
Replace Worker Nodes (Blue/Green Node Groups)
For worker nodes, we didn't immediately terminate the existing node group:
- We created a new managed node group with an EKS-optimized AMI compatible with the new Kubernetes version.
- Once the new nodes joined the cluster and showed
Ready, we started moving the workloads. - Cordon old node: Marked the node unschedulable so new pods were only scheduled on the new node group.
- Drain one node at a time:
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data. - We didn't drain everything together — we moved workloads gradually to maintain application availability.
How Did We Avoid Downtime?
Zero downtime was guaranteed because our critical microservices were engineered for High Availability:
Continuous Telemetry & Monitoring
Using CloudWatch and Prometheus/Grafana, continuously monitored throughout the upgrade:
- Node Health: Memory/CPU pressure and kubelet status.
- Pod Restarts: Monitored CrashLoopBackOff and restart counts.
- Pending Pods: Detected scheduling bottlenecks or resource exhaustion.
- CPU and Memory: Monitored cluster-wide headroom and OOM warnings.
- ALB Target Health: Confirmed targets remained healthy in AWS Target Groups.
- Application Latency: Verified p95/p99 response times did not degrade.
- 5xx HTTP Errors: Monitored error rates to confirm zero dropped requests.
Validate Before Removing Anything
Once all workloads were running on the new node group, we didn't immediately remove the old one:
- We performed smoke testing and validated critical application flows under real traffic.
- Once everything was verified stable, we cleanly deleted the old node group.
- Then we repeated the same process for subsequent minor version bumps (e.g. v1.34 → v1.35 → v1.36 → v1.37).
- Conducted rigorous pre-checks: EKS Upgrade Insights, API deprecations (Pluto), add-on compatibility, and lower-environment testing.
- Upgraded strictly one minor version at a time (v1.34 → v1.35 → v1.36 → v1.37): control plane first, followed by managed add-ons.
- Implemented blue/green worker node replacement with new EKS-optimized AMI node groups; cordoned and drained nodes one-by-one.
- Guaranteed zero downtime via HA safeguards: replica counts ≥ 2, PodDisruptionBudgets, readiness probes, preStop hooks, and multi-AZ spread.
- Monitored CloudWatch/Grafana telemetry (5xx errors, latency, pending pods, ALB target health) throughout.
- Executed critical business flow smoke tests before safely decommissioning old node groups.