⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Observability & Monitoring Interview Questions Scenario 88 of 90 in Observability & Monitoring
Senior DevOps / SRE Observability Prometheus & Metrics Core Concept

Q: What is the difference between a Gauge and a Counter in Prometheus? When should you use each, how do rate() and irate() work on counters, and what happens when a process restarts?

Mastering Gauge vs Counter in Prometheus: cumulative monotonically increasing counters vs fluctuating gauges, how rate() handles server restarts and resets, and when to choose each metric type.

#gauge vs counter prometheus #counter vs gauge metric #Prometheus #PromQL #Metrics #rate() #irate() #Observability #SRE
🎙️ Candidate Opening & Architectural Context
"In production telemetry, choosing the wrong Prometheus metric type between a Counter and a Gauge undermines all alerting and SLO accuracy. A Counter only increases (or resets to 0 on crash), whereas a Gauge can go up or down arbitrarily."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Fundamental Architectural Differences

Differentiating Counters and Gauges at the data collection level:

  • Counter: A cumulative metric that represents a single monotonically increasing counter. Its value can only increase or reset to zero upon process restart. Examples: http_requests_total, system_cpu_seconds_total, errors_total.
  • Gauge: A metric that represents a single numerical value that can arbitrarily go up and down. It captures a snapshot in time. Examples: node_memory_MemAvailable_bytes, jvm_threads_current, temperature_celsius, queue_depth.
Pro Tip: Never use a Gauge for counting events (like HTTP requests or errors) because Prometheus scrapes at intervals (e.g., every 15s) and any events occurring between scrape intervals will be permanently lost.
2️⃣

PromQL Query Functions & Counter Reset Compensation

How Prometheus mathematical operators interact with each type:

  • rate(v[range]): Calculates the per-second average rate of increase across a range vector. Critically, rate() automatically detects and compensates for counter resets (when a pod restarts and resets to 0).
  • irate(v[range]): Instant rate based on the last two data points in the range window. Ideal for fast-moving spikes, but sensitive to scrape jitter.
  • delta(v[range]) / deriv(v[range]): Used exclusively on Gauges to calculate differences or slopes. Never apply rate() to a Gauge because counter-reset compensation will corrupt the results when a gauge drops naturally!
Advertisement
3️⃣

Practical PromQL Query Patterns

Real-world queries used by SRE teams for alerts and dashboards:

  • Counter rule: Always pass a counter through rate() before applying aggregations like sum() or avg(). Running sum(http_requests_total) without rate causes artificial step jumps whenever pods restart.
  • Gauge rule: Can be directly aggregated with sum(queue_length) or avg(cpu_temperature).
4️⃣

Decision Matrix: Counter vs Gauge

Fast reference for instrumentation design:

  • Requests / Orders / Errors: Always Counter (use _total suffix per Prometheus naming conventions).
  • Memory / CPU % / Connections: Always Gauge.
  • Queue Size / In-Flight Requests: Gauge (track active requests via gauge increment on start, decrement on finish).
  • Duration / Latency: Histogram or Summary (which under the hood exposes a counter for _count and _sum).
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Counters measure rate of occurrence over time with automatic crash-reset compensation. Gauges measure instantaneous state at the moment of scrape. Never apply rate() to a Gauge, and never aggregate raw counters before calculating rate()."
⚡ 60-Second Elevator Pitch Talking Points
  • A Counter is a cumulative monotonically increasing metric used for countable events like requests and errors; it only goes up or resets to 0 on restart.
  • A Gauge represents snapshot values that fluctuate up and down arbitrarily, such as memory usage, disk space, or current active connections.
  • In PromQL, counters are queried with rate() to calculate per-second frequency while automatically smoothing over process restarts, whereas gauges use raw values or delta().
Advertisement
Want more Observability & Monitoring scenarios?
Explore our complete collection of scenario-based Observability & Monitoring interview runbooks.
Browse All Observability & Monitoring Questions →