Q: Prometheus is a pull-based system, meaning it scrapes targets that are continuously running. How do you monitor a cron job that runs for only 3 seconds and terminates before Prometheus has a chance to scrape it?
You use the Prometheus Pushgateway.
#Observability #Observability #L1 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Pushgateway architecture.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Production Solution & Architecture
You use the Prometheus Pushgateway. The Pushgateway is an intermediary, continuously running component. The short-lived cron job, right before it terminates, actively *pushes* its final metrics (like job_duration_seconds or items_processed) to the Pushgateway via an HTTP POST. The Pushgateway stores these metrics in memory indefinitely. Prometheus can then leisurely scrape the Pushgateway on its standard interval (e.g., every 15 seconds) to collect the metrics of the dead job.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: You use the Prometheus Pushgateway.."
⚡ 60-Second Elevator Pitch Talking Points
- Immediate Triage: You use the Prometheus Pushgateway.
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement