⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Observability & Monitoring Interview Questions Scenario 88 of 88 in Observability & Monitoring
Senior SRE / DevOps Observability Latency Analysis & Distributed Tracing Latency Deep Dive

Q: Users are reporting slowness, but CPU and memory look normal. What will you check next?

Troubleshooting application latency when CPU and RAM utilization metrics are comfortably low. Diagnosing thread pool exhaustion, DB connection queueing, disk iowait, and DNS timeouts.

#SRE #Performance #Latency #Tracing #Bottlenecks #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When CPU and Memory metrics look calm (e.g., 20% CPU, 30% RAM) but users experience severe latency or timeouts, the system is almost certainly trapped in a 'wait state'—requests are waiting in queues, blocked on locks, waiting for disk I/O, or stalled on external dependencies."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Analyze Distributed Tracing & Span Waterfall

Open APM distributed traces (OpenTelemetry, Jaeger, Datadog). Look at the trace waterfall: is time spent inside application processing, waiting on an external HTTP call, or waiting on a database query?

# Look for spans with high duration but low CPU time
2

Check Thread Pool & Connection Pool Queueing

If the application's HTTP worker thread pool (e.g., Tomcat, Gunicorn, Node.js event loop) or database connection pool (HikariCP, PgBouncer) is exhausted, incoming requests are queued. Queued requests accumulate latency without consuming any CPU.

# Check JVM thread states or connection pool active vs max metrics:
# Example: hikaricp_pending_threads > 0
# Example: nodejs_eventloop_lag_seconds > 0.1
3

Check Disk I/O Wait (wa) & Storage IOPS Saturation

If processes are waiting on slow disks or EBS volume IOPS limits, CPU utilization appears low because the CPU is idling waiting for data. Check the 'wa' (iowait) column in vmstat and %util in iostat.

vmstat 1 5
# Check the 'wa' column (I/O wait)
iostat -xz 1 5
# Check %util and await (>10ms indicates storage throttling)
4

Check Network Contention, TCP Retransmissions & DNS Delays

DNS timeouts (like the Linux 5-second glibc / CoreDNS race condition) or TCP packet loss cause requests to hang without CPU consumption. Check network metrics and packet drop counters.

netstat -s | grep -i retrans
# Test CoreDNS resolution latency:
dig +time=2 +tries=1 @<CoreDNS_IP> api.service.internal
5

Inspect Downstream Dependency & Lock Contention

Check if an external payment gateway, third-party authentication API, or database row-level locking (e.g. pg_stat_activity showing 'exclusive lock') is serializing requests.

# In PostgreSQL:
SELECT pid, state, wait_event_type, wait_event, query FROM pg_stat_activity WHERE wait_event IS NOT NULL;
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Low CPU and low Memory with high latency indicates a 'wait state'. Look at thread and connection pool queues, disk I/O wait, distributed trace spans, and downstream lock contention."
⚡ 60-Second Elevator Pitch Talking Points
  • Low CPU + High Latency is the classic symptom of a wait state: requests are blocked waiting for a shared resource.
  • Inspect APM distributed tracing first to pinpoint which downstream span (DB, Redis, 3rd party API) is holding the request.
  • Check thread pools and connection pool queues: if all connections are checked out, requests sit in a queue without burning CPU.
  • Run 'vmstat 1' and check 'wa' (iowait) and 'iostat -xz 1': storage IOPS throttling causes processes to sleep in 'D' state.
  • Verify database locks and network retransmissions: row-level locking or DNS 5-second packet drops stall traffic silently.
Advertisement
Want more Observability & Monitoring scenarios?
Explore our complete collection of scenario-based Observability & Monitoring interview runbooks.
Browse All Observability & Monitoring Questions →