⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Observability Production Scenario [L2]

Q: An application generates 500,000 log lines per second. Storing everything costs $100,000/month. You need detailed debugging capability but cannot afford full-volume storage. What sampling strategy allows you to capture errors while discarding routine logs?

Log sampling reduces volume while preserving critical signals. There are two approaches:

#Observability #Observability #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""During a high-traffic production event, our observability stack proved essential in isolating this latency surge. The interviewer is testing: Log cost optimization, sampling strategies, tail-based sampling.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Log sampling reduces volume while preserving critical signals. There are two approaches:

  • All logs from *error* requests (even if they're low-volume)
  • All logs from *slow* requests (latency > 1000ms)
  • Random sample of *success* requests (1% to track healthy profiles)
  • Full volume: 500,000 logs/sec × $10 per million logs = $150,000/month
  • With tail-based sampling (errors + 1% sample):
2️⃣

Remediation & Permanent Safeguards

1. Head-Based Sampling (Probabilistic): At the point where the log is generated, randomly decide: keep this log with probability P (e.g., 1% of logs). The decision is made instantly, lowest CPU overhead. Problem: You randomly discard errors. With 1% sampling, you'll miss 99% of the stack traces. 2. Tail-Based Sampling (Intelligent): Capture *all* logs in a temporary buffer, but only ship to the aggregator if they match certain criteria: A log forwarder like Fluentbit or OpenTelemetry Collector buffers logs in memory as they arrive, tags them with request outcome (error/success/latency), and makes shipping decisions *after* the request completes. Example (with OpenTelemetry): Cost Math:

import random
if random.random() < 0.01:  # Keep 1%
    logger.info("request completed")
  • Errors: ~5,000/sec, success sample: ~5,000/sec = 10,000 logs/sec
  • Monthly cost: 10,000 × 86400 × 30 × $10 / 1,000,000 = $2,592/month (98% savings)
  • Debugging capability: All errors captured + representative success traces for normal behavior analysis.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: All logs from *error* requests (even if they're low-volume)."
⚡ 60-Second Elevator Pitch Talking Points
  • All logs from *error* requests (even if they're low-volume)
  • All logs from *slow* requests (latency > 1000ms)
  • Random sample of *success* requests (1% to track healthy profiles)
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability