Q: Your observability infrastructure costs $200,000/month (Datadog, Prometheus, etc.), but the CFO demands a 40% cost reduction. You cannot lose visibility into production. Design a cost-aware observability strategy with specific architectural changes.
Cost reduction requires architectural restructuring—not just turning off features. The strategy is Hot/Warm/Cold tiers with intelligent r...
#Observability #Observability #L3 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Business-aware engineering, observability architecture, cost vs reliability tradeoffs.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Cost reduction requires architectural restructuring—not just turning off features. The strategy is Hot/Warm/Cold tiers with intelligent routing.
- Full-resolution metrics (15-second granularity) for 30 days in Datadog = billions of data points @ $0.05 per 1000
- All logs ingested into Splunk/Datadog for instant searchability = $$$
- Metrics Tiering:
- Hot (3 days): Full-resolution (15s) Prometheus on cheap local hardware. Covers "right now" incident response. Cost: ~$1,000/month (hardware).
- Warm (7-30 days): Downsampled (5-minute resolution) shipped to S3/Thanos with query-on-demand. Cost: ~$500/month (storage).
- Cold (>30 days): Parquet/ORC format in S3. Queries require Athena (serverless) scanning. Cost: ~$100/month (occasional audits).
- Savings: From $80k/month in Datadog to $1.6k/month.
- Log Tiering:
- Hot (7 days): High-priority logs only (errors, warnings) in Elasticsearch. Cost: ~$2,000/month.
- Warm (30 days): All logs (unindexed) in S3 gzip archives. Query via Athena/Splunk on-demand. Cost: ~$500/month.
- Cold (>30 days): Compliance/audit archives (immutable, rarely retrieved). Cost: ~$50/month.
- Savings: From $90k/month in Datadog logs to $2.55k/month.
- Traces (formerly 100% sampled at all times):
- Intelligent Sampling: Jaeger/Datadog samples at 0.5% baseline (random), but bumps to 100% for:
- Errors (always trace failures)
- High latency requests (>500ms)
- Specific high-value transactions (payment checkout)
2️⃣
Remediation & Permanent Safeguards
Current High-Cost Architecture: Cost-Optimized Architecture: Operational Changes: Total Monthly Cost Reduction: Trade-offs Accepted:
- Savings: From $30k/month in trace storage to $3k/month (only interesting requests traced).
- Alerts & Notification:
- Move from expensive alerting (Datadog Monitors @ $40/monitor) to cheaper open-source tools:
- Prometheus AlertManager + custom webhook integrations (free)
- Grafana alerts (open-source, self-hosted) (free)
- Savings: $40k/month in monitor licensing to ~$500/month infrastructure.
- On-call engineers know: "For the last 30 days, query the hot Elasticsearch. For 30-day-old issues, run Athena queries (5-min query latency)."
- For post-mortems, sacrifice instant query time, enable Athena scanning of S3 (acceptable, not urgent).
- Batch jobs and non-critical services use only logs + metrics (no traces).
- Before: Datadog basic tier at $200k/month
- After: Self-hosted Prometheus/Grafana/ELK + S3 = $8.5k/month
- Savings: 95.75% cost reduction ($191.5k/month)
- Instant instant query latency lost (but alerts still fast)
- Team must relearn on-call procedures
- Requires in-house expertise to maintain ELK and Prometheus
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Full-resolution metrics (15-second granularity) for 30 days in Datadog = billions of data points @ $0.05 per 1000."
⚡ 60-Second Elevator Pitch Talking Points
- Full-resolution metrics (15-second granularity) for 30 days in Datadog = billions of data points ...
- All logs ingested into Splunk/Datadog for instant searchability = $$$
- Metrics Tiering:
Advertisement