⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] General DevOps General DevOps — Scenario-Based Interview Questions Staff SRE Scenario [L3]

Q: Your auto-scaling group scales out based purely on CPU utilization crossing 80%. An intensive marketing email blast goes out at exactly 9:00 AM, causing traffic to spike 10x instantly. The servers crash before new instances can boot. How do you handle this?

Threshold-based CPU scaling is Reactive. By the time CPU hits 80%, the application is already stressed. Booting EC2 instances takes minut...

#General DevOps #General DevOps — Scenario-Based Interview Questions #L3 #DevOps #SRE #Architecture
🎙️ Candidate Opening & Architectural Context
""In our engineering organization, DevOps culture meant aligning developer speed with site reliability. The interviewer is testing: Reactive vs Proactive scaling, predictive scaling, scheduled actions.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Threshold-based CPU scaling is Reactive. By the time CPU hits 80%, the application is already stressed. Booting EC2 instances takes minutes—by the time they are ready, the initial servers have already collapsed under the instant 10x spike.

  • Scheduled Scaling: Apply an ASG Scheduled Action to artificially inflate the minimum capacity from 2 to 20 instances at 8:45 AM, giving them 15 minutes to comfortably boot and register before the 9:00 AM blast.
  • Predictive Scaling: Use AWS Predictive Scaling, which uses Machine Learning to analyze historical daily/weekly traffic patterns and pre-warms capacity automatically before anticipated spikes occur.
2️⃣

Remediation & Permanent Safeguards

To handle known traffic spikes, you must use Proactive Scaling:

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Scheduled Scaling: Apply an ASG Scheduled Action to artificially inflate the minimum capacity from 2 to 20 instances at 8:45 AM, g."
⚡ 60-Second Elevator Pitch Talking Points
  • Scheduled Scaling: Apply an ASG Scheduled Action to artificially inflate the minimum capacity fro...
  • Predictive Scaling: Use AWS Predictive Scaling, which uses Machine Learning to analyze historical...
Advertisement
Want more General DevOps scenarios?
Explore our complete collection of scenario-based General DevOps interview runbooks.
Browse All General DevOps Questions →

📚 Related Production Scenarios in General DevOps