Q: When measuring API latency, why is an Average (Mean) a terrible metric compared to Percentiles (P95, P99)?
An Average hides extreme outliers. If 99 users experience lightning-fast 10ms latencies, but 1 user hits a database timeout and waits 5,0...
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
An Average hides extreme outliers. If 99 users experience lightning-fast 10ms latencies, but 1 user hits a database timeout and waits 5,000ms, the mathematical average is ~60ms. It looks perfectly healthy on a dashboard, masking the fact that a user had a terrible, broken experience. Percentiles (like P99) order all requests from fastest to slowest. A P99 of 800ms means that 99% of requests were faster than 800ms, and the worst 1% of users experienced 800ms or worse. Alerting on P99 or P99.9 ensures you are monitoring the "long-tail" latency, protecting the experience of your most heavily impacted customers rather than just the majority.
- Immediate Triage: An Average hides extreme outliers. If 99 users experience lightning-fast 10ms latencies, but 1
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.