Q: What is the meaning of P50 and P90 latency? Explain P50 vs P90 vs P99 latency percentiles, why average latency is dangerously misleading in production, and how to measure percentiles using Prometheus histograms.
Explaining P50, P90, P95, and P99 latency percentiles: why arithmetic mean averages lie in distributed systems, understanding multimodal distributions, and calculating histogram_quantile() in Prometheus.
Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Definitions: P50, P90, P95, and P99
Understanding percentile rankings:
- P50 (50th Percentile / Median): 50% of all requests were served faster than this duration, and 50% were slower. It represents the typical experience of an ordinary user.
- P90 (90th Percentile): 90% of requests completed faster than this time, while the slowest 10% took longer. It catches early degradation before the majority of users feel it.
- P99 (99th Percentile / Tail Latency): 99% of requests completed faster, representing the worst 1% of user experiences. In high-traffic systems (10,000 req/sec), 1% means 100 requests every second are suffering this slow latency!
Why Arithmetic Mean (Average) Lies
The fatal flaw of averages in production:
- Imagine 100 requests: 95 requests finish in 10ms (cache hits), but 5 requests stall in a database deadlock and take 10,000ms (10 seconds).
- Average:
(95 * 10 + 5 * 10000) / 100 = 509.5 ms. The average looks like a mild 500ms delay, completely hiding the fact that 5% of users were blocked for 10 full seconds! - Percentiles: P50 is 10ms. P90 is 10ms. But P95 and P99 are 10,000ms! Percentiles immediately highlight the long-tail degradation.
Calculating Percentiles with Prometheus Histograms
How to query P50, P90, and P99 in PromQL:
- Always configure histogram buckets with exponential increments spanning from fast responses (e.g. 5ms) to timeouts (e.g. 10s).
Diagnosing Tail Latency (P90 / P99 Spikes)
Common root causes when P50 is fast but P99 is slow:
- Garbage Collection Pauses: JVM/Go stop-the-world GC pauses pausing execution periodically.
- Connection Pool Exhaustion: 90% of requests find an idle DB connection; 10% must wait for a connection to release.
- Resource Contention / CPU Throttling: CFS quota throttling on Kubernetes pods with rigid CPU limits.
- Cold Cache Misses: Infrequently accessed database keys requiring disk reads.
- P50 latency is the median where 50% of requests are faster, while P90 and P99 measure the threshold where 90% and 99% of requests complete faster.
- We never use average latency in production because a small percentage of 10-second timeouts is mathematically diluted by fast cache hits, hiding catastrophic failures.
- In SRE, we track P95 and P99 using Prometheus histogram_quantile() to alert on connection pool starvation, GC pauses, and CPU throttling before users complain.