Q: Users are reporting slowness, but CPU and memory look normal. What will you check next?
Troubleshooting application latency when CPU and RAM utilization metrics are comfortably low. Diagnosing thread pool exhaustion, DB connection queueing, disk iowait, and DNS timeouts.
Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Analyze Distributed Tracing & Span Waterfall
Open APM distributed traces (OpenTelemetry, Jaeger, Datadog). Look at the trace waterfall: is time spent inside application processing, waiting on an external HTTP call, or waiting on a database query?
# Look for spans with high duration but low CPU time
Check Thread Pool & Connection Pool Queueing
If the application's HTTP worker thread pool (e.g., Tomcat, Gunicorn, Node.js event loop) or database connection pool (HikariCP, PgBouncer) is exhausted, incoming requests are queued. Queued requests accumulate latency without consuming any CPU.
# Check JVM thread states or connection pool active vs max metrics:
# Example: hikaricp_pending_threads > 0
# Example: nodejs_eventloop_lag_seconds > 0.1
Check Disk I/O Wait (wa) & Storage IOPS Saturation
If processes are waiting on slow disks or EBS volume IOPS limits, CPU utilization appears low because the CPU is idling waiting for data. Check the 'wa' (iowait) column in vmstat and %util in iostat.
vmstat 1 5
# Check the 'wa' column (I/O wait)
iostat -xz 1 5
# Check %util and await (>10ms indicates storage throttling)
Check Network Contention, TCP Retransmissions & DNS Delays
DNS timeouts (like the Linux 5-second glibc / CoreDNS race condition) or TCP packet loss cause requests to hang without CPU consumption. Check network metrics and packet drop counters.
netstat -s | grep -i retrans
# Test CoreDNS resolution latency:
dig +time=2 +tries=1 @<CoreDNS_IP> api.service.internal
Inspect Downstream Dependency & Lock Contention
Check if an external payment gateway, third-party authentication API, or database row-level locking (e.g. pg_stat_activity showing 'exclusive lock') is serializing requests.
# In PostgreSQL:
SELECT pid, state, wait_event_type, wait_event, query FROM pg_stat_activity WHERE wait_event IS NOT NULL;
- Low CPU + High Latency is the classic symptom of a wait state: requests are blocked waiting for a shared resource.
- Inspect APM distributed tracing first to pinpoint which downstream span (DB, Redis, 3rd party API) is holding the request.
- Check thread pools and connection pool queues: if all connections are checked out, requests sit in a queue without burning CPU.
- Run 'vmstat 1' and check 'wa' (iowait) and 'iostat -xz 1': storage IOPS throttling causes processes to sleep in 'D' state.
- Verify database locks and network retransmissions: row-level locking or DNS 5-second packet drops stall traffic silently.