⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 57 of 98 in FinOps & System Design
Staff SRE / Observability Architect System Design Observability & Tracing Architecture System Design

Q: Your microservices fleet generates 1,000,000 trace spans per second. Ingesting and storing 100% of traces into an APM backend would cost $180,000/month and crash storage nodes. However, random head-based sampling drops the exact traces containing 5xx HTTP errors and high-latency anomalies that SREs need during incidents. How do you design an OpenTelemetry tail-based sampling pipeline that captures 100% of errors and slow requests while discarding 95% of healthy traces?

Engineering a high-throughput distributed tracing platform processing 1,000,000 spans/sec using OpenTelemetry Collector tail-based sampling, load-balancing exporters, and Grafana Tempo object storage tiering.

#System Design #Distributed Tracing #OpenTelemetry #Tail Sampling #Tempo #Kafka
🎙️ Candidate Opening & Architectural Context
"Head-based sampling decides whether to drop a trace at the initial ingress request before knowing whether the downstream database query will fail or take 10 seconds. We architected a two-tier OpenTelemetry collector cluster with Load-Balancing exporters and Tail-Based Sampling."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Architect Two-Tier OpenTelemetry Collector Architecture

Ensure all spans belonging to the exact same TraceID land on the identical collector instance:

  • Tier 1 (Agent Collectors): Deployed lightweight OTel Collectors as DaemonSets receiving spans locally from application SDKs via OTLP/gRPC.
  • Load-Balancing Exporter: Tier 1 uses loadbalancingexporter with traceID consistent hashing (routing_key: trace_id) to route spans to Tier 2 collectors.
  • Deterministic Routing: All spans for trace 4bf92f3577... across 50 different microservices are guaranteed to route to the exact same Tier 2 collector pod.
Pro Tip: Tail-based sampling requires assembling the entire trace in memory; consistent hashing by traceID across Tier 1 collectors makes this possible without central distributed locking.
2️⃣

Configure Intelligent Tail-Based Sampling Policies

Evaluate the complete trace lifecycle before deciding whether to persist to storage:

  • Decision Window: Configured decision_wait: 10s, holding spans in an in-memory buffer until all child spans arrive.
  • Policy 1 - Always Keep Errors: status_code: [ERROR] or any span attribute containing http.status_code >= 500 -> Keep 100%.
  • Policy 2 - Keep Slow Traces: Any trace with total duration > 1,500ms -> Keep 100%.
  • Policy 3 - Probabilistic Baseline: All remaining healthy HTTP 200 fast traces -> Keep 1% for baseline latency distributions.
Pro Tip: Tail-based sampling guarantees that every single production incident, exception, or timeout is captured with 100% fidelity.
3️⃣

Enforce Memory Limiter Processor & Circuit Breaking

Prevent OOMKilled crashes during massive incident error storms:

  • Memory Limiter Processor: Configured memory_limiter with check_interval: 1s, limit_percentage: 80, and spike_limit_percentage: 20.
  • Graceful Degradation: If in-memory buffers approach 80% RAM during an error tsunami, the collector automatically falls back to probabilistic head-sampling to protect collector stability.
Pro Tip: Without a memory limiter, a massive cascade of errors will overwhelm collector memory, taking down the entire observability infrastructure.
4️⃣

Export to Grafana Tempo & Measure Cost Reductions

Persist sampled spans to object-storage backed distributed trace storage:

  • Storage Engine: Exported filtered spans to Grafana Tempo backed by AWS S3 / Google Cloud Storage.
  • Ingestion Reduction: Slashed active ingested spans from 1,000,000 spans/sec to 58,000 spans/sec (94.2% data reduction).
  • Cost Impact: Monthly tracing storage and compute bills plummeted from $180,000/month to $8,400/month while increasing incident debuggability.
Pro Tip: Storing traces in S3 with Parquet/block compression allows querying traces directly from Grafana dashboards without paying search index taxes.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Tail-based tracing with two-tier OpenTelemetry collectors and consistent traceID hashing allows buffering requests in memory to retain 100% of errors and latency outliers while dropping 95% of healthy traces."
⚡ 60-Second Elevator Pitch Talking Points
  • Route spans to Tier 2 collectors using traceID consistent hashing in the load-balancing exporter.
  • Buffer traces for 10 seconds to evaluate full transaction health before making sampling decisions.
  • Retain 100% of HTTP 5xx errors and slow traces (>1.5s) while sampling healthy traces at 1%.
  • Cut storage ingestion fees by over 90% while guaranteeing zero missed production incident traces.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →