Q: EC2 CPU suddenly reaches 100% — how would you troubleshoot?
Triage runbook for handling sudden 100% CPU utilization on production EC2 instances: telemetry triage, identifying culprit processes, handling legit vs malicious load, and permanent safeguards.
#AWS #EC2 #CloudWatch #Linux #Troubleshooting #Incident Triage
🎙️ Candidate Opening & Architectural Context
"When a production EC2 instance hits 100% CPU, my priority is two-fold: stop customer impact immediately (mitigation) while capturing telemetry to pinpoint the exact culprit process (root cause analysis)."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Triage From Outside (CloudWatch & Telemetry)
Before logging in, look at multi-dimensional metrics to classify the nature of the spike:
- Correlate Metrics: In CloudWatch, compare
CPUUtilizationagainstNetworkIn,NetworkOut, andDiskReadOps/DiskWriteBytes. - Traffic Surge vs Rogue Process: If high CPU coincides with a 10x spike in NetworkIn and ALB requests, it is likely legitimate or DDoS traffic. If NetworkIn is flat but CPU spiked abruptly, it is likely an internal process, runaway job, or compromised instance.
- Check Auto Scaling Group (ASG): Check if ASG target tracking is scaling out instances or if the ASG has hit its
max_sizelimit.
2️⃣
SSH / SSM Session & Process Inspection
Log into the instance (using AWS Systems Manager Session Manager or SSH) and inspect active processes:
uptime: Check load average against total core count (nproc). If 4 cores and load is 40, system is severely saturated.top -corhtop: PressPto sort by CPU consumption. Note the PID, USER, and COMMAND.ps aux --sort=-%cpu | head -10: Print the top 10 CPU-consuming processes with full command-line arguments.vmstat 1 5: Check if CPU is spending time in user space (us), system/kernel (sy), or I/O wait (wa). High I/O wait indicates a storage bottleneck rather than pure computation.
3️⃣
Diagnose the Culprit Process
Examine the identified process to determine what it is doing:
- Thread Inspection: Run
top -H -p <PID>to inspect individual threads inside the process. - Syscall Tracing: Run
strace -p <PID> -cfor 10 seconds to identify spinning loops or repetitive failing syscalls. - Application Thread Dump: For Java/Node/Go runtimes, capture a thread dump (e.g.
jstack <PID>or Go pprof) before terminating. - Check Recent Crons: Check
/var/log/syslogor/var/log/cronfor heavy scheduled backup, antivirus, or indexing tasks.
4️⃣
Mitigate and Restore Service
Take decisive, safe actions based on the diagnosis:
- If Rogue Application Process: Issue
kill -15 <PID>(SIGTERM) for graceful shutdown. Only usekill -9if unresponsive. - If Legitimate Traffic Surge: Manually increase ASG desired capacity, or provision a read replica / cache if backend DB queries are choking the app.
- If Compromised / Crypto-Miner: Immediately isolate instance: change Security Group to quarantine SG (no outbound/inbound except security audit), capture EBS snapshot for forensics, and terminate.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never reboot an instance blindly. Capture thread dumps and top process telemetry first so you don't lose the smoking gun, and ensure auto-scaling and CPU alarms prevent single-node exhaustion."
⚡ 60-Second Elevator Pitch Talking Points
- Check CloudWatch: correlate CPU with NetworkIn and ALB traffic to differentiate traffic surge from rogue process.
- Access instance via SSM Session Manager / SSH; run 'top -c', 'uptime', and 'ps aux --sort=-%cpu | head -10'.
- Check 'vmstat 1 5' for 'us' (app code) vs 'wa' (I/O wait bottleneck) vs 'sy' (kernel context switching).
- Inspect threads with 'top -H -p <PID>' and capture thread dump (jstack/pprof) or strace before killing.
- Mitigate: graceful SIGTERM, scale out ASG, or quarantine if compromised.
- Prevent: CloudWatch Alarm at 75% CPU, ASG auto-scaling, CPU limits (cgroups/Docker), and query optimization.
Advertisement