⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] AWS ECS, EKS, Lambda Production Scenario [L2]

Q: Your ECS task keeps stopping with exit code 137. What's happening?

Exit code 137 = container killed by OOM (Out of Memory) — 128 + 9 (SIGKILL).

#AWS #ECS, EKS, Lambda #L2 #Cloud #Infrastructure
🎙️ Candidate Opening & Architectural Context
""In a previous role, our monitoring paged me for a similar incident across our AWS VPC infrastructure. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Exit code 137 = container killed by OOM (Out of Memory) — 128 + 9 (SIGKILL).

  • Check task definition memory hard/soft limit.
  • Increase the memory limit.
  • Or find and fix the memory leak in your application.
2️⃣

Remediation & Permanent Safeguards

The container exceeded its memory limit and was killed. Fix:

  • Use CloudWatch Container Insights to monitor actual memory usage trends.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Check task definition memory hard/soft limit.."
⚡ 60-Second Elevator Pitch Talking Points
  • Check task definition memory hard/soft limit.
  • Increase the memory limit.
  • Or find and fix the memory leak in your application.
Advertisement
Want more AWS scenarios?
Explore our complete collection of scenario-based AWS interview runbooks.
Browse All AWS Questions →

📚 Related Production Scenarios in AWS