Q: Prometheus graphs have regular gaps every 15 minutes for many targets. What do you investigate?
Regular gaps point to scrape or ingestion problems, not random application behavior.
#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Scrape reliability, timeouts, sample limits, target health.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Regular gaps point to scrape or ingestion problems, not random application behavior.
upfor affected targets during the gaps.scrape_duration_secondsto see whether scrapes are timing out.scrape_samples_scrapedand sample-limit errors if exporters produce too many metrics.- Prometheus CPU, memory, WAL, and disk I/O during the gap.
2️⃣
Remediation & Permanent Safeguards
I would check: The fix depends on the cause: increase scrape timeout carefully, reduce exporter work, shard Prometheus, or remove expensive metrics.
- Network, DNS, service discovery, or load balancer behavior on a 15-minute schedule.
- Exporter logs for slow collection, especially exporters that call cloud APIs or databases.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: up for affected targets during the gaps.."
⚡ 60-Second Elevator Pitch Talking Points
- up for affected targets during the gaps.
- scrape_duration_seconds to see whether scrapes are timing out.
- scrape_samples_scraped and sample-limit errors if exporters produce too many metrics.
Advertisement