Q: A Prometheus graph shows gaps for a service after pods restart. Does a missing line mean the service was healthy with zero traffic? How do you alert on missing data correctly?
Missing data does not mean zero. It usually means Prometheus did not scrape the target, the metric disappeared, or the pod restarted and ...
#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Our SRE team tackled this monitoring and metrics bottleneck to eliminate false-positive alert fatigue. The interviewer is testing: Prometheus staleness, absent metrics, scrape health.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Missing data does not mean zero. It usually means Prometheus did not scrape the target, the metric disappeared, or the pod restarted and the series went stale.
- Zero value: The service exported the metric with value
0. - No series: Prometheus has no current sample for that label set.
2️⃣
Remediation & Permanent Safeguards
I would separate two cases: For scrape health, alert on: For a metric that must always exist, use: or alert when expected targets disappear from service discovery. This prevents treating missing telemetry as healthy behavior.
up{job="checkout"} == 0
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Zero value: The service exported the metric with value 0.."
⚡ 60-Second Elevator Pitch Talking Points
- Zero value: The service exported the metric with value 0.
- No series: Prometheus has no current sample for that label set.
Advertisement