Q: You need an alert for P95 HTTP latency from Prometheus histogram metrics. What query shape do you use, and what mistake should you avoid?
For a classic Prometheus histogram, I would calculate P95 from the bucket rate:
#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Our SRE team tackled this monitoring and metrics bottleneck to eliminate false-positive alert fatigue. The interviewer is testing: Correct PromQL for histograms and percentile aggregation.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Production Solution & Architecture
For a classic Prometheus histogram, I would calculate P95 from the bucket rate: This says: calculate the 95th percentile latency per service over the last 5 minutes and alert if it is above 500ms. The major mistake is averaging per-pod P95 values: That is mathematically wrong because percentiles are not additive. A pod with 10 requests and a pod with 100,000 requests should not have equal weight. Aggregate buckets first, then calculate the percentile.
histogram_quantile(
0.95,
sum by (le, service) (
rate(http_request_duration_seconds_bucket[5m])
)
) > 0.5
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: For a classic Prometheus histogram, I would calculate P95 from the bucket rate:."
⚡ 60-Second Elevator Pitch Talking Points
- Immediate Triage: For a classic Prometheus histogram, I would calculate P95 from the bucket rate:
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement