Q: What is the difference between a Push-based monitoring system (like DataDog/StatsD) and a Pull-based system (like Prometheus)?
1. Push: The application (or an agent on the host) writes metrics actively and sends them over the network to a centralized aggregator en...
#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Architecture, network topologies, auto-discovery.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
*Pros:* Easier to span NATs or firewalls (outbound is usually allowed), great for ephemeral/serverless functions that die too fast to be scraped.
- Push: The application (or an agent on the host) writes metrics actively and sends them over the network to a centralized aggregator endpoint (e.g., Datadog, InfluxDB).
- Pull (Prometheus): The centralized server uses an HTTP GET request (scrape) to pull a
/metricsendpoint exposed by the application.
2️⃣
Remediation & Permanent Safeguards
*Cons:* Can overwhelm the central server with UDP floods, and the aggregator doesn't inherently know if an agent died vs simply has no data to send. *Pros:* The server controls the ingestion rate, preventing DDOS. It implicitly knows when a service is dead because the HTTP GET fails (up == 0). It heavily relies on Service Discovery (like Consul or Kubernetes API) to find targets dynamically.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Push: The application (or an agent on the host) writes metrics actively and sends them over the network to a centralized aggregato."
⚡ 60-Second Elevator Pitch Talking Points
- Push: The application (or an agent on the host) writes metrics actively and sends them over the n...
- Pull (Prometheus): The centralized server uses an HTTP GET request (scrape) to pull a /metrics en...
Advertisement