Q: Your cluster upgrade from 1.26 to 1.27 failed halfway through. Control plane is on 1.27 but worker nodes are still on 1.26. Is this okay?
Yes — this is a supported temporary state during upgrades. Kubernetes supports N-2 version skew between control plane and nodes. A 1.27 c...
#Kubernetes #Advanced Scenarios #L2 #Container Orchestration #K8s #Terraform State
🎙️ Candidate Opening & Architectural Context
""In one of our high-traffic production clusters, our SRE team handled this incident using a standardized runbook. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Yes — this is a supported temporary state during upgrades. Kubernetes supports N-2 version skew between control plane and nodes. A 1.27 control plane can manage 1.25, 1.26, and 1.27 nodes.
- Verify the control plane is healthy:
kubectl get nodes— control plane nodes should show 1.27. - Continue upgrading worker nodes one by one: drain, upgrade kubelet/kubectl/kubeadm, uncordon.
- Do not skip more than one minor version during upgrades.
2️⃣
Remediation & Permanent Safeguards
Next steps: Never upgrade worker nodes before the control plane — that would be an unsupported configuration.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Verify the control plane is healthy: kubectl get nodes — control plane nodes should show 1.27.."
⚡ 60-Second Elevator Pitch Talking Points
- Verify the control plane is healthy: kubectl get nodes — control plane nodes should show 1.27.
- Continue upgrading worker nodes one by one: drain, upgrade kubelet/kubectl/kubeadm, uncordon.
- Do not skip more than one minor version during upgrades.
Advertisement