Q: You are building a multi-tenant SaaS application. You need to segregate metrics so each enterprise customer can view their own latency. Why is adding a `tenant_id` label to every Prometheus metric a bad idea, and what should you do instead?
Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrophic Cardinality Explosion.
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrophic Cardinality Explosion. If you have 10,000 tenants, and each interacts with 50 endpoints across 5 HTTP methods and 4 status codes, multiplying these combinations creates tens of millions of distinct metric series, which will quickly crash Prometheus due to OOM errors or bankrupt you in Datadog custom metric billing. Instead: Fast, aggregated system health (Metrics) should *not* be split by customer. To provide per-tenant dashboards, you should inject the tenant_id exclusively into Logs or Distributed Traces. Those systems are built to index high-cardinality metadata cheaply. You can then use tools like Datadog Log Analytics or Elasticsearch to graph latency specifically filtered by tenant_id without breaking the core metric TSDB.
- Immediate Triage: Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrop
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.