⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE AWS Compute & EC2 Production Incident

Q: EC2 CPU suddenly reaches 100% — how would you troubleshoot?

Triage runbook for handling sudden 100% CPU utilization on production EC2 instances: telemetry triage, identifying culprit processes, handling legit vs malicious load, and permanent safeguards.

#AWS #EC2 #CloudWatch #Linux #Troubleshooting #Incident Triage
🎙️ Candidate Opening & Architectural Context
"When a production EC2 instance hits 100% CPU, my priority is two-fold: stop customer impact immediately (mitigation) while capturing telemetry to pinpoint the exact culprit process (root cause analysis)."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Triage From Outside (CloudWatch & Telemetry)

Before logging in, look at multi-dimensional metrics to classify the nature of the spike:

  • Correlate Metrics: In CloudWatch, compare CPUUtilization against NetworkIn, NetworkOut, and DiskReadOps/DiskWriteBytes.
  • Traffic Surge vs Rogue Process: If high CPU coincides with a 10x spike in NetworkIn and ALB requests, it is likely legitimate or DDoS traffic. If NetworkIn is flat but CPU spiked abruptly, it is likely an internal process, runaway job, or compromised instance.
  • Check Auto Scaling Group (ASG): Check if ASG target tracking is scaling out instances or if the ASG has hit its max_size limit.
2️⃣

SSH / SSM Session & Process Inspection

Log into the instance (using AWS Systems Manager Session Manager or SSH) and inspect active processes:

  • uptime: Check load average against total core count (nproc). If 4 cores and load is 40, system is severely saturated.
  • top -c or htop: Press P to sort by CPU consumption. Note the PID, USER, and COMMAND.
  • ps aux --sort=-%cpu | head -10: Print the top 10 CPU-consuming processes with full command-line arguments.
  • vmstat 1 5: Check if CPU is spending time in user space (us), system/kernel (sy), or I/O wait (wa). High I/O wait indicates a storage bottleneck rather than pure computation.
3️⃣

Diagnose the Culprit Process

Examine the identified process to determine what it is doing:

  • Thread Inspection: Run top -H -p <PID> to inspect individual threads inside the process.
  • Syscall Tracing: Run strace -p <PID> -c for 10 seconds to identify spinning loops or repetitive failing syscalls.
  • Application Thread Dump: For Java/Node/Go runtimes, capture a thread dump (e.g. jstack <PID> or Go pprof) before terminating.
  • Check Recent Crons: Check /var/log/syslog or /var/log/cron for heavy scheduled backup, antivirus, or indexing tasks.
4️⃣

Mitigate and Restore Service

Take decisive, safe actions based on the diagnosis:

  • If Rogue Application Process: Issue kill -15 <PID> (SIGTERM) for graceful shutdown. Only use kill -9 if unresponsive.
  • If Legitimate Traffic Surge: Manually increase ASG desired capacity, or provision a read replica / cache if backend DB queries are choking the app.
  • If Compromised / Crypto-Miner: Immediately isolate instance: change Security Group to quarantine SG (no outbound/inbound except security audit), capture EBS snapshot for forensics, and terminate.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never reboot an instance blindly. Capture thread dumps and top process telemetry first so you don't lose the smoking gun, and ensure auto-scaling and CPU alarms prevent single-node exhaustion."
⚡ 60-Second Elevator Pitch Talking Points
  • Check CloudWatch: correlate CPU with NetworkIn and ALB traffic to differentiate traffic surge from rogue process.
  • Access instance via SSM Session Manager / SSH; run 'top -c', 'uptime', and 'ps aux --sort=-%cpu | head -10'.
  • Check 'vmstat 1 5' for 'us' (app code) vs 'wa' (I/O wait bottleneck) vs 'sy' (kernel context switching).
  • Inspect threads with 'top -H -p <PID>' and capture thread dump (jstack/pprof) or strace before killing.
  • Mitigate: graceful SIGTERM, scale out ASG, or quarantine if compromised.
  • Prevent: CloudWatch Alarm at 75% CPU, ASG auto-scaling, CPU limits (cgroups/Docker), and query optimization.
Advertisement
Want more AWS scenarios?
Explore our complete collection of scenario-based AWS interview runbooks.
Browse All AWS Questions →

📚 Related Production Scenarios in AWS