Q: You want to get an alert when your EC2 instance CPU exceeds 80% for more than 5 minutes. How do you set this up?
1. In CloudWatch, create an Alarm:
#AWS #Monitoring & CloudWatch #L2 #Cloud #Infrastructure #EC2
🎙️ Candidate Opening & Architectural Context
""AWS reliability requires differentiating between AWS control plane limits and host-level resource exhaustion. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Or do it all via CLI:
- In CloudWatch, create an Alarm:
- Metric:
EC2 → Per-Instance Metrics → CPUUtilization - Statistic: Average
- Period: 5 minutes (300 seconds)
- Condition:
> 80
2️⃣
Remediation & Permanent Safeguards
aws cloudwatch put-metric-alarm \
--alarm-name HighCPU \
--metric-name CPUUtilization \
--namespace AWS/EC2 \
--statistic Average \
--period 300 \
--threshold 80 \
--comparison-operator GreaterThanThreshold \
--evaluation-periods 1 \
--alarm-actions arn:aws:sns:...
- Evaluation periods: 1 (alarm triggers after 1 period = 5 minutes above threshold)
- Add an SNS action — send notification to an SNS topic.
- Subscribe your email or PagerDuty endpoint to the SNS topic.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: In CloudWatch, create an Alarm:."
⚡ 60-Second Elevator Pitch Talking Points
- In CloudWatch, create an Alarm:
- Metric: EC2 → Per-Instance Metrics → CPUUtilization
- Statistic: Average
Advertisement