Q: Assume an EC2 instance becomes intermittently unresponsive between 2 PM and 5 PM IST. When you check disk, CPU, and memory utilization, everything appears normal, but the system queue length is increasing. How would you troubleshoot and resolve this issue?
Deep troubleshooting runbook for an EC2 instance that becomes unresponsive during specific afternoon hours with normal CPU/memory metrics but rising disk queue length.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Correlate with EBS IOPS & Throughput Burst Balance
Check Amazon CloudWatch metrics for the attached EBS volumes (`VolumeQueueLength`, `BurstBalance`, `VolumeReadOps`, `VolumeWriteOps`). If using `gp2` volumes, burst credits may deplete precisely between 2 PM and 5 PM due to an afternoon batch process, throttling throughput down to baseline and causing I/O requests to queue up indefinitely.
aws cloudwatch get-metric-statistics \
--namespace AWS/EBS \
--metric-name BurstBalance \
--dimensions Name=VolumeId,Value=vol-0123456789abcdef0 \
--start-time 2026-10-06T08:00:00Z \
--end-time 2026-10-06T12:00:00Z \
--period 300 \
--statistics Average
Inspect Instance-Level EBS Bandwidth Throttling
Every EC2 instance type (e.g. `t3.medium`, `m5.large`) has a maximum EBS-optimized bandwidth limit separate from the volume limit. If the instance hits `EBSIOBalance%` or `EBSByteBalance%` limits in CloudWatch, disk operations stall even if the gp3 volume itself is within limits.
# On the Linux EC2 instance:
iostat -xz 1 10
# High await (e.g. >50ms) and high %util with low CPU indicates disk bottleneck
Identify Uninterruptible Sleep (D-State) Processes
Run `ps -eo state,pid,cmd | grep '^D'` during the 2 PM - 5 PM window. This displays processes waiting on kernel disk locks, NFS mounts, or network sockets.
ps -eo state,pid,user,cmd | awk '$1 ~ /D/ {print $0}'
# Inspect what file the process is blocked on:
cat /proc/<pid>/stack
Migrate to gp3 / Provisioned IOPS & Upgrade Instance
Upgrade the volume from `gp2` to `gp3` (which provides dedicated baseline 3,000 IOPS and 125 MB/s with no burst balance depletion) and resize the EC2 instance to a higher EBS-optimized bandwidth tier.
- Check CloudWatch VolumeQueueLength and BurstBalance metrics to confirm EBS I/O throttling.
- Run iostat -xz and inspect D-state processes (ps -eo state,pid,cmd) on the Linux host.
- Verify instance-level EBS-optimized bandwidth saturation.
- Migrate from gp2 to gp3 to obtain consistent baseline IOPS and throughput without credit exhaustion.