⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Observability Procedure #1: Clear Deadlock Production Scenario [L2]

Q: After deploying OpenTelemetry instrumentation, traces appear in staging but not production. The application logs show spans are being created. Where do you look first?

If spans are created inside the application but do not reach the backend, the problem is usually between the SDK and the telemetry backend.

#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""Logs tell you what happened, metrics tell you where to look, and distributed traces pinpoint the exact slow component. The interviewer is testing: OpenTelemetry pipeline debugging, collector/exporter configuration.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

If spans are created inside the application but do not reach the backend, the problem is usually between the SDK and the telemetry backend.

  • Collector endpoint: Confirm production points to the correct OpenTelemetry Collector address and protocol (grpc vs http/protobuf).
  • Collector pipelines: Verify the traces pipeline has a receiver, required processors, and the correct exporter wired together.
  • Exporter errors: Look at collector logs and metrics such as send failures, queue size, dropped spans, and retry counts.
2️⃣

Remediation & Permanent Safeguards

I would check: The fastest test is to enable collector debug logging or temporarily export to a local logging exporter. That proves whether spans reached the collector before debugging the vendor/backend side.

  • Network and auth: Check NetworkPolicy, security groups, proxy settings, API keys, and TLS certificates.
  • Sampling: Confirm production is not configured with an accidental 0% sampler.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Collector endpoint: Confirm production points to the correct OpenTelemetry Collector address and protocol (grpc vs http/protobuf).."
⚡ 60-Second Elevator Pitch Talking Points
  • Collector endpoint: Confirm production points to the correct OpenTelemetry Collector address and ...
  • Collector pipelines: Verify the traces pipeline has a receiver, required processors, and the corr...
  • Exporter errors: Look at collector logs and metrics such as send failures, queue size, dropped sp...
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability