Q: A Terraform change wants to replace a production EKS node group, but the cluster has critical workloads. How do you approach it?
Avoid a blind replacement. Create a new node group with the desired configuration, allow nodes to join, drain workloads gradually with re...
#Terraform #Use VPC ID from another module #L3 #IaC #Cloud Infrastructure
🎙️ Candidate Opening & Architectural Context
""When terraform plan shows unexpected changes, my golden rule is: never apply blindly. Investigate the diff first. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Production Solution & Architecture
Avoid a blind replacement. Create a new node group with the desired configuration, allow nodes to join, drain workloads gradually with respect for PodDisruptionBudgets, and then remove the old node group after capacity is healthy. Terraform can manage both node groups during the transition. This reduces risk compared with letting one resource replacement decide the whole rollout.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Avoid a blind replacement. Create a new node group with the desired configuration, allow nodes to join, drain workloads gradually ."
⚡ 60-Second Elevator Pitch Talking Points
- Immediate Triage: Avoid a blind replacement. Create a new node group with the desired configuration, allow nodes
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement