⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Observability Production Scenario [L2]

Q: You are building a multi-tenant SaaS application. You need to segregate metrics so each enterprise customer can view their own latency. Why is adding a `tenant_id` label to every Prometheus metric a bad idea, and what should you do instead?

Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrophic Cardinality Explosion.

#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Cardinality explosions, log vs metrics cost structures.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Production Solution & Architecture

Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrophic Cardinality Explosion. If you have 10,000 tenants, and each interacts with 50 endpoints across 5 HTTP methods and 4 status codes, multiplying these combinations creates tens of millions of distinct metric series, which will quickly crash Prometheus due to OOM errors or bankrupt you in Datadog custom metric billing. Instead: Fast, aggregated system health (Metrics) should *not* be split by customer. To provide per-tenant dashboards, you should inject the tenant_id exclusively into Logs or Distributed Traces. Those systems are built to index high-cardinality metadata cheaply. You can then use tools like Datadog Log Analytics or Elasticsearch to graph latency specifically filtered by tenant_id without breaking the core metric TSDB.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrophic Cardinality Explosion.."
⚡ 60-Second Elevator Pitch Talking Points
  • Immediate Triage: Adding a tenant_id label to Prometheus metrics is a fatal mistake because it causes a catastrop
  • Run targeted verification commands before modifying configuration.
  • Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability