⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Observability Staff SRE Scenario [L3]

Q: Your observability infrastructure costs $200,000/month (Datadog, Prometheus, etc.), but the CFO demands a 40% cost reduction. You cannot lose visibility into production. Design a cost-aware observability strategy with specific architectural changes.

Cost reduction requires architectural restructuring—not just turning off features. The strategy is Hot/Warm/Cold tiers with intelligent r...

#Observability #Observability #L3 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Business-aware engineering, observability architecture, cost vs reliability tradeoffs.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Cost reduction requires architectural restructuring—not just turning off features. The strategy is Hot/Warm/Cold tiers with intelligent routing.

  • Full-resolution metrics (15-second granularity) for 30 days in Datadog = billions of data points @ $0.05 per 1000
  • All logs ingested into Splunk/Datadog for instant searchability = $$$
  • Metrics Tiering:
  • Hot (3 days): Full-resolution (15s) Prometheus on cheap local hardware. Covers "right now" incident response. Cost: ~$1,000/month (hardware).
  • Warm (7-30 days): Downsampled (5-minute resolution) shipped to S3/Thanos with query-on-demand. Cost: ~$500/month (storage).
  • Cold (>30 days): Parquet/ORC format in S3. Queries require Athena (serverless) scanning. Cost: ~$100/month (occasional audits).
  • Savings: From $80k/month in Datadog to $1.6k/month.
  • Log Tiering:
  • Hot (7 days): High-priority logs only (errors, warnings) in Elasticsearch. Cost: ~$2,000/month.
  • Warm (30 days): All logs (unindexed) in S3 gzip archives. Query via Athena/Splunk on-demand. Cost: ~$500/month.
  • Cold (>30 days): Compliance/audit archives (immutable, rarely retrieved). Cost: ~$50/month.
  • Savings: From $90k/month in Datadog logs to $2.55k/month.
  • Traces (formerly 100% sampled at all times):
  • Intelligent Sampling: Jaeger/Datadog samples at 0.5% baseline (random), but bumps to 100% for:
  • Errors (always trace failures)
  • High latency requests (>500ms)
  • Specific high-value transactions (payment checkout)
2️⃣

Remediation & Permanent Safeguards

Current High-Cost Architecture: Cost-Optimized Architecture: Operational Changes: Total Monthly Cost Reduction: Trade-offs Accepted:

  • Savings: From $30k/month in trace storage to $3k/month (only interesting requests traced).
  • Alerts & Notification:
  • Move from expensive alerting (Datadog Monitors @ $40/monitor) to cheaper open-source tools:
  • Prometheus AlertManager + custom webhook integrations (free)
  • Grafana alerts (open-source, self-hosted) (free)
  • Savings: $40k/month in monitor licensing to ~$500/month infrastructure.
  • On-call engineers know: "For the last 30 days, query the hot Elasticsearch. For 30-day-old issues, run Athena queries (5-min query latency)."
  • For post-mortems, sacrifice instant query time, enable Athena scanning of S3 (acceptable, not urgent).
  • Batch jobs and non-critical services use only logs + metrics (no traces).
  • Before: Datadog basic tier at $200k/month
  • After: Self-hosted Prometheus/Grafana/ELK + S3 = $8.5k/month
  • Savings: 95.75% cost reduction ($191.5k/month)
  • Instant instant query latency lost (but alerts still fast)
  • Team must relearn on-call procedures
  • Requires in-house expertise to maintain ELK and Prometheus
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Full-resolution metrics (15-second granularity) for 30 days in Datadog = billions of data points @ $0.05 per 1000."
⚡ 60-Second Elevator Pitch Talking Points
  • Full-resolution metrics (15-second granularity) for 30 days in Datadog = billions of data points ...
  • All logs ingested into Splunk/Datadog for instant searchability = $$$
  • Metrics Tiering:
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability