⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Observability Procedure #1: Clear Deadlock Staff SRE Scenario [L3]

Q: During an incident, dashboards show "no data" for several critical services. How do you distinguish a telemetry outage from an application outage?

Observability systems need their own monitoring.

#Observability #Procedure #1: Clear Deadlock #L3 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""In an interview, I explain how we designed actionable, symptom-based alerting using the Four Golden Signals. The interviewer is testing: Meta-monitoring and observability reliability.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Observability systems need their own monitoring.

  • Collector and log shipper health: dropped data, queue size, exporter failures, restarts.
  • Prometheus scrape health and target discovery.
  • Remote-write or vendor ingestion status.
  • Independent blackbox checks against the application.
2️⃣

Remediation & Permanent Safeguards

I would check: If blackbox checks and platform metrics are healthy but telemetry pipelines are failing, it is a monitoring incident. If both user-facing probes and telemetry are bad, it is likely an application or infrastructure incident. I would also alert separately on telemetry pipeline failure, because "no data" during an outage is itself a severe operational risk.

  • Cloud/load balancer metrics that do not depend on the same telemetry pipeline.
  • Recent deploys or config changes to collectors and agents.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Collector and log shipper health: dropped data, queue size, exporter failures, restarts.."
⚡ 60-Second Elevator Pitch Talking Points
  • Collector and log shipper health: dropped data, queue size, exporter failures, restarts.
  • Prometheus scrape health and target discovery.
  • Remote-write or vendor ingestion status.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability