Q: An SRE says the infrastructure is stable, the dev team says the system is slow, and monitoring dashboards show everything is GREEN. Who do you believe — and what do you check first?
How to resolve the classic production paradox where infrastructure dashboards show all green, yet developers and users complain of acute system slowness.
Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Replace Average (p50) Metrics with High-Percentile Latency (p95, p99, p99.9)
Averages hide catastrophic outages. If 900 users complete requests in 20ms and 100 users time out at 30,000ms, the mathematical average is ~3 seconds (which might look 'acceptable' on a coarse dashboard), but 10% of users are completely blocked. Query p95 and p99 latency immediately.
# PromQL query to expose tail latency
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, path))
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, path))
Analyze Edge-to-Backend Latency Discrepancy
Compare metrics recorded at the client edge (CDN / CloudFront / ALB) against metrics recorded inside the backend application pods. If ALB latency is 5 seconds but application code reports 50ms execution time, the delay is happening in the queuing network layer: TLS handshake overhead, reverse-proxy backlog drops, or slow client transfer.
Inspect Resource Saturation vs. Resource Exhaustion (USE Method)
Green CPU doesn't mean healthy compute. Check for CFS CPU Throttling (`container_cpu_cfs_throttled_periods_total`) where Kubernetes throttles a container due to strict CPU limits even when the node has 80% free CPU. Check database connection pool wait queues and thread lock contention.
# PromQL: Detect hidden CPU throttling on Kubernetes pods
sum(rate(container_cpu_cfs_throttled_periods_total[5m])) by (pod)
/
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod) * 100
Trace End-to-End Distributed Tracing Spans (OpenTelemetry)
Pull up Jaeger/Tempo distributed traces for affected requests. Look at the waterfall visualization: identify whether time is spent waiting on PostgreSQL row locks, DNS resolution, Redis socket connect timeouts, or third-party API gateways.
- Believe the users and developers immediately: host averages mask severe tail-latency outages.
- Switch dashboards from average latency to p95 and p99 percentiles to expose degradation.
- Check for CFS CPU Throttling where containers are artificially choked despite low node CPU.
- Use OpenTelemetry distributed traces to pinpoint the exact microservice span or database lock causing delays.