⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Observability Procedure #1: Clear Deadlock Staff SRE Scenario [L3]

Q: After enabling service mesh telemetry, Prometheus cardinality explodes because metrics include source pod, destination pod, path, method, response code, and workload labels. How do you control it?

Service mesh telemetry is powerful but can create a series for every source-destination-path combination.

#Observability #Procedure #1: Clear Deadlock #L3 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: Service mesh metrics, label control, aggregation strategy.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Service mesh telemetry is powerful but can create a series for every source-destination-path combination.

  • Dropping pod-level labels from high-volume metrics and keeping workload, namespace, and service labels.
  • Normalizing paths, such as /orders/{id} instead of /orders/12345.
  • Keeping method and response-code class, but avoiding unnecessary headers or user-level labels.
  • Creating recording rules for common service-to-service RED metrics.
2️⃣

Remediation & Permanent Safeguards

I would control it by: The goal is service-level observability, not a unique time series for every request shape.

  • Applying metric relabeling at scrape time to remove labels that are not used in alerts or dashboards.
  • Setting cardinality budgets per team and reviewing top series regularly.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Dropping pod-level labels from high-volume metrics and keeping workload, namespace, and service labels.."
⚡ 60-Second Elevator Pitch Talking Points
  • Dropping pod-level labels from high-volume metrics and keeping workload, namespace, and service l...
  • Normalizing paths, such as /orders/{id} instead of /orders/12345.
  • Keeping method and response-code class, but avoiding unnecessary headers or user-level labels.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability