⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Observability & Monitoring Interview Questions Scenario 91 of 96 in Observability & Monitoring
Staff SRE / Principal Platform Engineer Observability SRE Golden Signals & Percentiles SRE Philosophy

Q: An SRE says the infrastructure is stable, the dev team says the system is slow, and monitoring dashboards show everything is GREEN. Who do you believe — and what do you check first?

How to resolve the classic production paradox where infrastructure dashboards show all green, yet developers and users complain of acute system slowness.

#Observability #SRE #p99 Latency #APM #Golden Signals #SLO #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"I always believe the users and the developers. Traditional infrastructure dashboards only monitor host-level resource utilization (CPU < 50%, Memory < 60%, Disk < 40%) and average latency (p50). A system can have 10% CPU usage while 95% of users experience 10-second timeouts due to lock contention, DNS latency, thread pool starvation, or tail latency (p99/p99.9) hidden by averages."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Replace Average (p50) Metrics with High-Percentile Latency (p95, p99, p99.9)

Averages hide catastrophic outages. If 900 users complete requests in 20ms and 100 users time out at 30,000ms, the mathematical average is ~3 seconds (which might look 'acceptable' on a coarse dashboard), but 10% of users are completely blocked. Query p95 and p99 latency immediately.

# PromQL query to expose tail latency
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, path))
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, path))
2

Analyze Edge-to-Backend Latency Discrepancy

Compare metrics recorded at the client edge (CDN / CloudFront / ALB) against metrics recorded inside the backend application pods. If ALB latency is 5 seconds but application code reports 50ms execution time, the delay is happening in the queuing network layer: TLS handshake overhead, reverse-proxy backlog drops, or slow client transfer.

Pro Tip: SRE Axiom: Never trust internal pod metrics without verifying client-facing Synthetic Monitors and edge access logs.
Advertisement
3

Inspect Resource Saturation vs. Resource Exhaustion (USE Method)

Green CPU doesn't mean healthy compute. Check for CFS CPU Throttling (`container_cpu_cfs_throttled_periods_total`) where Kubernetes throttles a container due to strict CPU limits even when the node has 80% free CPU. Check database connection pool wait queues and thread lock contention.

# PromQL: Detect hidden CPU throttling on Kubernetes pods
sum(rate(container_cpu_cfs_throttled_periods_total[5m])) by (pod) 
/ 
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod) * 100
4

Trace End-to-End Distributed Tracing Spans (OpenTelemetry)

Pull up Jaeger/Tempo distributed traces for affected requests. Look at the waterfall visualization: identify whether time is spent waiting on PostgreSQL row locks, DNS resolution, Redis socket connect timeouts, or third-party API gateways.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Always believe user complaints over green infrastructure graphs. Green graphs usually mean servers aren't crashing, not that requests are succeeding fast. Inspect p99 tail latency, CPU CFS throttling, and distributed trace waterfalls."
⚡ 60-Second Elevator Pitch Talking Points
  • Believe the users and developers immediately: host averages mask severe tail-latency outages.
  • Switch dashboards from average latency to p95 and p99 percentiles to expose degradation.
  • Check for CFS CPU Throttling where containers are artificially choked despite low node CPU.
  • Use OpenTelemetry distributed traces to pinpoint the exact microservice span or database lock causing delays.
Advertisement
Want more Observability & Monitoring scenarios?
Explore our complete collection of scenario-based Observability & Monitoring interview runbooks.
Browse All Observability & Monitoring Questions →