⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Observability Procedure #1: Clear Deadlock Production Scenario [L2]

Q: After a Kubernetes upgrade, Prometheus shows `up == 0` for many pod scrape targets. What do you check?

I would troubleshoot the scrape path from Prometheus to the pods.

#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""During a high-traffic production event, our observability stack proved essential in isolating this latency surge. The interviewer is testing: Kubernetes service discovery, scraping, network and auth issues.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

I would troubleshoot the scrape path from Prometheus to the pods.

  • Service discovery: Are the pods still discovered with the expected labels and annotations?
  • Endpoint changes: Did ServiceMonitor, PodMonitor, or scrape configs stop matching after label changes?
  • NetworkPolicy: Can Prometheus still reach pod IPs and metrics ports?
  • TLS/auth: Did certificates, service account tokens, or mTLS settings change?
2️⃣

Remediation & Permanent Safeguards

up == 0 is a scrape failure. The application may be healthy, but Prometheus cannot collect its metrics.

  • Metrics endpoint: Does /metrics still respond from inside the cluster?
  • Prometheus logs: Look for scrape errors such as timeout, connection refused, 401, 403, or x509 failures.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Service discovery: Are the pods still discovered with the expected labels and annotations?."
⚡ 60-Second Elevator Pitch Talking Points
  • Service discovery: Are the pods still discovered with the expected labels and annotations?
  • Endpoint changes: Did ServiceMonitor, PodMonitor, or scrape configs stop matching after label cha...
  • NetworkPolicy: Can Prometheus still reach pod IPs and metrics ports?
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability