Q: An application generates 500,000 log lines per second. Storing everything costs $100,000/month. You need detailed debugging capability but cannot afford full-volume storage. What sampling strategy allows you to capture errors while discarding routine logs?
Log sampling reduces volume while preserving critical signals. There are two approaches:
🛠️ Production Runbook & Step-by-Step Resolution
Initial Diagnostics & Root Cause Analysis
Log sampling reduces volume while preserving critical signals. There are two approaches:
- All logs from *error* requests (even if they're low-volume)
- All logs from *slow* requests (latency > 1000ms)
- Random sample of *success* requests (1% to track healthy profiles)
- Full volume: 500,000 logs/sec × $10 per million logs = $150,000/month
- With tail-based sampling (errors + 1% sample):
Remediation & Permanent Safeguards
1. Head-Based Sampling (Probabilistic): At the point where the log is generated, randomly decide: keep this log with probability P (e.g., 1% of logs). The decision is made instantly, lowest CPU overhead. Problem: You randomly discard errors. With 1% sampling, you'll miss 99% of the stack traces. 2. Tail-Based Sampling (Intelligent): Capture *all* logs in a temporary buffer, but only ship to the aggregator if they match certain criteria: A log forwarder like Fluentbit or OpenTelemetry Collector buffers logs in memory as they arrive, tags them with request outcome (error/success/latency), and makes shipping decisions *after* the request completes. Example (with OpenTelemetry): Cost Math:
import random
if random.random() < 0.01: # Keep 1%
logger.info("request completed")
- Errors: ~5,000/sec, success sample: ~5,000/sec = 10,000 logs/sec
- Monthly cost: 10,000 × 86400 × 30 × $10 / 1,000,000 = $2,592/month (98% savings)
- Debugging capability: All errors captured + representative success traces for normal behavior analysis.
- All logs from *error* requests (even if they're low-volume)
- All logs from *slow* requests (latency > 1000ms)
- Random sample of *success* requests (1% to track healthy profiles)