Q: Explain the difference between `Gauge` and `Counter` metric types in Prometheus.
1. Counter: A cumulative metric that can only go up (or reset to zero on restart). Examples include http_requests_total or bytes_sent. Be...
#Observability #Observability #L1 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Fundamental metric types.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
- Counter: A cumulative metric that can *only go up* (or reset to zero on restart). Examples include
http_requests_totalorbytes_sent. Because it only goes up, you never query its raw value directly; you always apply a rate function (e.g.,rate(http_requests_total[5m])) to see how fast it's growing. - Gauge: A metric that can arbitrarily *go up and down* over time. Examples include
cpu_memory_usage,current_queue_depth, ortemperature. You can query gauges directly to evaluate their current value without needing a rate function.
2️⃣
Remediation & Permanent Safeguards
Execute the resolution runbook and verify workload health:
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Counter: A cumulative metric that can *only go up* (or reset to zero on restart). Examples include http_requests_total or bytes_se."
⚡ 60-Second Elevator Pitch Talking Points
- Counter: A cumulative metric that can *only go up* (or reset to zero on restart). Examples includ...
- Gauge: A metric that can arbitrarily *go up and down* over time. Examples include cpu_memory_usag...
Advertisement