Q: Your EC2 instance is showing high CPU and your application is slow. What steps do you take?
1. CloudWatch metrics — check CPU utilization over time. Is it sustained or spiky?
#AWS #EC2 & Compute #L2 #Cloud #Infrastructure #EC2
🎙️ Candidate Opening & Architectural Context
""When an interviewer asks how I troubleshoot this in AWS, I frame it through my hands-on production experience. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
- CloudWatch metrics — check CPU utilization over time. Is it sustained or spiky?
- SSH in and check —
toporhtopto see which process is consuming CPU. - Check for runaway processes — a stuck process, an infinite loop, a cron job misfiring.
- Check memory — sometimes high CPU is actually a memory issue causing constant swap. Check
free -mandvmstat.
2️⃣
Remediation & Permanent Safeguards
Execute the resolution runbook and verify workload health:
- Scale vertically — if it's a legitimate load, stop the instance, change instance type to a larger one, restart.
- Scale horizontally — add it to an Auto Scaling Group, put a Load Balancer in front, scale out.
- Right-sizing — use AWS Compute Optimizer recommendations.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: CloudWatch metrics — check CPU utilization over time. Is it sustained or spiky?."
⚡ 60-Second Elevator Pitch Talking Points
- CloudWatch metrics — check CPU utilization over time. Is it sustained or spiky?
- SSH in and check — top or htop to see which process is consuming CPU.
- Check for runaway processes — a stuck process, an infinite loop, a cron job misfiring.
Advertisement