⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] AWS EC2 & Compute Staff SRE Scenario [L3]

Q: An EC2 instance in an Auto Scaling Group keeps being terminated and replaced. The new instance starts, becomes healthy, then gets terminated again in a cycle. What's happening?

This sounds like a lifecycle hook or health check issue:

#AWS #EC2 & Compute #L3 #Cloud #Infrastructure #ALB
🎙️ Candidate Opening & Architectural Context
""AWS reliability requires differentiating between AWS control plane limits and host-level resource exhaustion. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

This sounds like a lifecycle hook or health check issue:

  • Application actually is unhealthy — the new instance really is failing. Check the health check endpoint. Maybe there's a deployment bug.
  • ALB health check failing immediately — instance isn't ready to serve traffic before health checks kick in. The ASG uses ALB health checks and terminates the instance before the app is fully started.
  • Lifecycle hook stuck — a lifecycle hook (autoscaling:EC2_INSTANCE_LAUNCHING) is not completing. Instance is in Pending:Wait state but something is killing it.
2️⃣

Remediation & Permanent Safeguards

Check ASG activity logs in the console — it will say exactly why each instance was terminated.

  • Scale-in protection not set — instance is being terminated by scale-in event despite looking fine.
  • EC2 instance reachability check failing — hardware issue at the hypervisor level. Check AWS console system status checks.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Application actually is unhealthy — the new instance really is failing. Check the health check endpoint. Maybe there's a deploymen."
⚡ 60-Second Elevator Pitch Talking Points
  • Application actually is unhealthy — the new instance really is failing. Check the health check en...
  • ALB health check failing immediately — instance isn't ready to serve traffic before health checks...
  • Lifecycle hook stuck — a lifecycle hook (autoscaling:EC2_INSTANCE_LAUNCHING) is not completing. I...
Advertisement
Want more AWS scenarios?
Explore our complete collection of scenario-based AWS interview runbooks.
Browse All AWS Questions →

📚 Related Production Scenarios in AWS