Q: Your Prometheus Time-Series Database (TSDB) is running on a massive disk with plenty of space left, but it is thrashing the CPU and IOPS with high "Compaction" activity. What causes excessive compaction?
High TSDB compaction (and resulting IOPS thrashing) is heavily correlated with Metric Churn.
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
High TSDB compaction (and resulting IOPS thrashing) is heavily correlated with Metric Churn. Churn is different from High Cardinality. Churn happens when a service creates brand new metric series, stops updating them after highly ephemeral periods, and creates new ones. For example, if a developer mistakenly uses the Kubernetes Pod IP as a metric label in a rapidly auto-scaling environment. Every time a pod is replaced, the old metric series is abandoned, and a new one is created. Prometheus groups recent data in temporary memory blocks. When moving to persistent disk, it "compacts" related series. Massive churn forces the compactor to constantly rewrite indices and stitch together millions of fragmented, short-lived series, burning massive CPU. The fix is to remove ephemeral labels (like Pod IPs) and use static identifiers (like Service Names).
- Immediate Triage: High TSDB compaction (and resulting IOPS thrashing) is heavily correlated with Metric Churn.
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.