⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AWS & Cloud Architecture Interview Questions Scenario 175 of 177 in AWS & Cloud Architecture
Senior Cloud Engineer (L2) AWS EC2 Performance & EBS Throttling L2 Cloud Screen

Q: Assume an EC2 instance becomes intermittently unresponsive between 2 PM and 5 PM IST. When you check disk, CPU, and memory utilization, everything appears normal, but the system queue length is increasing. How would you troubleshoot and resolve this issue?

Deep troubleshooting runbook for an EC2 instance that becomes unresponsive during specific afternoon hours with normal CPU/memory metrics but rising disk queue length.

#AWS #EC2 #EBS #IOPS #Disk Queue #CloudWatch #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When an EC2 instance experiences increasing system queue length (run queue or disk I/O queue) while CPU and RAM utilization appear normal, the operating system threads are blocked in an uninterruptible sleep state (`D` state in Linux). This indicates that processes are waiting on storage I/O, EBS burst balance exhaustion, network bandwidth throttling, or hypervisor CPU credit exhaustion."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Correlate with EBS IOPS & Throughput Burst Balance

Check Amazon CloudWatch metrics for the attached EBS volumes (`VolumeQueueLength`, `BurstBalance`, `VolumeReadOps`, `VolumeWriteOps`). If using `gp2` volumes, burst credits may deplete precisely between 2 PM and 5 PM due to an afternoon batch process, throttling throughput down to baseline and causing I/O requests to queue up indefinitely.

aws cloudwatch get-metric-statistics \
  --namespace AWS/EBS \
  --metric-name BurstBalance \
  --dimensions Name=VolumeId,Value=vol-0123456789abcdef0 \
  --start-time 2026-10-06T08:00:00Z \
  --end-time 2026-10-06T12:00:00Z \
  --period 300 \
  --statistics Average
2

Inspect Instance-Level EBS Bandwidth Throttling

Every EC2 instance type (e.g. `t3.medium`, `m5.large`) has a maximum EBS-optimized bandwidth limit separate from the volume limit. If the instance hits `EBSIOBalance%` or `EBSByteBalance%` limits in CloudWatch, disk operations stall even if the gp3 volume itself is within limits.

# On the Linux EC2 instance:
iostat -xz 1 10
# High await (e.g. >50ms) and high %util with low CPU indicates disk bottleneck
Advertisement
3

Identify Uninterruptible Sleep (D-State) Processes

Run `ps -eo state,pid,cmd | grep '^D'` during the 2 PM - 5 PM window. This displays processes waiting on kernel disk locks, NFS mounts, or network sockets.

ps -eo state,pid,user,cmd | awk '$1 ~ /D/ {print $0}'
# Inspect what file the process is blocked on:
cat /proc/<pid>/stack
4

Migrate to gp3 / Provisioned IOPS & Upgrade Instance

Upgrade the volume from `gp2` to `gp3` (which provides dedicated baseline 3,000 IOPS and 125 MB/s with no burst balance depletion) and resize the EC2 instance to a higher EBS-optimized bandwidth tier.

Pro Tip: Production Fix: Migrate legacy gp2 volumes to gp3 immediately. gp3 provides guaranteed baseline IOPS without relying on burst credit balances.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Normal CPU and RAM with rising queue length indicates threads blocked in uninterruptible disk I/O sleep (D state). Check EBS BurstBalance and instance-level EBS bandwidth throttling, and migrate to gp3."
⚡ 60-Second Elevator Pitch Talking Points
  • Check CloudWatch VolumeQueueLength and BurstBalance metrics to confirm EBS I/O throttling.
  • Run iostat -xz and inspect D-state processes (ps -eo state,pid,cmd) on the Linux host.
  • Verify instance-level EBS-optimized bandwidth saturation.
  • Migrate from gp2 to gp3 to obtain consistent baseline IOPS and throughput without credit exhaustion.
Advertisement
Want more AWS & Cloud Architecture scenarios?
Explore our complete collection of scenario-based AWS & Cloud Architecture interview runbooks.
Browse All AWS & Cloud Architecture Questions →