Q: What is the difference between a Gauge and a Counter in Prometheus? When should you use each, how do rate() and irate() work on counters, and what happens when a process restarts?
Mastering Gauge vs Counter in Prometheus: cumulative monotonically increasing counters vs fluctuating gauges, how rate() handles server restarts and resets, and when to choose each metric type.
Want to master this scenario in a live sandbox? The Linux Foundation's Prometheus Certified Associate (PCA) & Monitoring Labs covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Fundamental Architectural Differences
Differentiating Counters and Gauges at the data collection level:
- Counter: A cumulative metric that represents a single monotonically increasing counter. Its value can only increase or reset to zero upon process restart. Examples:
http_requests_total,system_cpu_seconds_total,errors_total. - Gauge: A metric that represents a single numerical value that can arbitrarily go up and down. It captures a snapshot in time. Examples:
node_memory_MemAvailable_bytes,jvm_threads_current,temperature_celsius,queue_depth.
PromQL Query Functions & Counter Reset Compensation
How Prometheus mathematical operators interact with each type:
- rate(v[range]): Calculates the per-second average rate of increase across a range vector. Critically,
rate()automatically detects and compensates for counter resets (when a pod restarts and resets to 0). - irate(v[range]): Instant rate based on the last two data points in the range window. Ideal for fast-moving spikes, but sensitive to scrape jitter.
- delta(v[range]) / deriv(v[range]): Used exclusively on Gauges to calculate differences or slopes. Never apply
rate()to a Gauge because counter-reset compensation will corrupt the results when a gauge drops naturally!
Practical PromQL Query Patterns
Real-world queries used by SRE teams for alerts and dashboards:
- Counter rule: Always pass a counter through
rate()before applying aggregations likesum()oravg(). Runningsum(http_requests_total)without rate causes artificial step jumps whenever pods restart. - Gauge rule: Can be directly aggregated with
sum(queue_length)oravg(cpu_temperature).
Decision Matrix: Counter vs Gauge
Fast reference for instrumentation design:
- Requests / Orders / Errors: Always Counter (use
_totalsuffix per Prometheus naming conventions). - Memory / CPU % / Connections: Always Gauge.
- Queue Size / In-Flight Requests: Gauge (track active requests via gauge increment on start, decrement on finish).
- Duration / Latency: Histogram or Summary (which under the hood exposes a counter for
_countand_sum).
- A Counter is a cumulative monotonically increasing metric used for countable events like requests and errors; it only goes up or resets to 0 on restart.
- A Gauge represents snapshot values that fluctuate up and down arbitrarily, such as memory usage, disk space, or current active connections.
- In PromQL, counters are queried with rate() to calculate per-second frequency while automatically smoothing over process restarts, whereas gauges use raw values or delta().