Q: A tracing backend becomes expensive and slow because span names include full URLs like `/users/123/orders/987`. What is the problem and how do you fix it?
The span name contains unbounded identifiers. Every user ID and order ID creates a different operation name, making search, aggregation, ...
#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Our SRE team tackled this monitoring and metrics bottleneck to eliminate false-positive alert fatigue. The interviewer is testing: Span naming, high-cardinality trace attributes.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Production Solution & Architecture
The span name contains unbounded identifiers. Every user ID and order ID creates a different operation name, making search, aggregation, and storage expensive. I would normalize span names: Then I would store IDs only as attributes if they are safe and necessary, and avoid indexing high-cardinality attributes by default. Good span names represent the operation shape, not one specific request. This lets the tracing backend group latency, errors, and throughput by endpoint correctly.
Bad: GET /users/123/orders/987
Good: GET /users/{user_id}/orders/{order_id}
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: The span name contains unbounded identifiers. Every user ID and order ID creates a different operation name, making search, aggreg."
⚡ 60-Second Elevator Pitch Talking Points
- Immediate Triage: The span name contains unbounded identifiers. Every user ID and order ID creates a different op
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement