⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Observability Staff SRE Scenario [L3]

Q: Your Prometheus instance is crashing with OOM every few hours. You identify the culprit: a Kubernetes deployment metric with a `pod_name` label containing every pod UUID ever created in the cluster (including deleted pods). Why is cardinality so deadly, and how do you prevent this without restarting Prometheus?

Each unique combination of label values creates a separate time-series. Prometheus stores each series' metadata, recent data points, and ...

#Observability #Observability #L3 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Our SRE team tackled this monitoring and metrics bottleneck to eliminate false-positive alert fatigue. The interviewer is testing: Understanding metric cardinality, label design, active remediation.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Each unique combination of label values creates a separate time-series. Prometheus stores each series' metadata, recent data points, and indices in memory.

  • 10 jobs × 100 instances × 50,000 pod UUIDs × 10 containers = 500 million series
  • Each series occupies ~1KB of memory (metadata, indices) = 500GB needed (will OOM instantly on a 64GB server).
  • Label Design (Prevent): Never use unbounded identifiers like pod_name or user_id as labels. Use only bounded dimensions:
2️⃣

Remediation & Permanent Safeguards

The Math: If you have a metric with labels {job, instance, pod_name, container}: Why deletes matter: When a pod is deleted, Prometheus doesn't immediately purge its cardinality. The metric remains "recorded" until the TSDB compaction cycle runs (days later), so cardinality grows unbounded. Prevention & Remediation: Reload Prometheus: kill -HUP (no restart, configs reloaded live). New scraped metrics won't have pod_name, and Prometheus will garbage-collect the old series over time. If a metric's cardinality exceeds 10,000, auto-alert before it crashes.

# Bad: pod_name has infinite cardinality
   http_requests_total{method, status, pod_name}
   
   # Good: namespace and deployment are bounded (~100s)
   http_requests_total{method, status, namespace, deployment}
  • Metric Relabeling (Active Remediation): Without restarting, add metric_relabel_configs to your scrape config to drop the problematic label:
  • Cardinality Budgets: Proactively monitor cardinality:
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: 10 jobs × 100 instances × 50,000 pod UUIDs × 10 containers = 500 million series."
⚡ 60-Second Elevator Pitch Talking Points
  • 10 jobs × 100 instances × 50,000 pod UUIDs × 10 containers = 500 million series
  • Each series occupies ~1KB of memory (metadata, indices) = 500GB needed (will OOM instantly on a 6...
  • Label Design (Prevent): Never use unbounded identifiers like pod_name or user_id as labels. Use o...
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability