⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE [L2] Observability Procedure #1: Clear Deadlock Production Scenario [L2]

Q: Every alert in your system has severity `critical`, so on-call gets paged for low-risk issues. How do you design alert severity levels?

Severity should map to required human action.

#Observability #Procedure #1: Clear Deadlock #L2 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""During a high-traffic production event, our observability stack proved essential in isolating this latency surge. The interviewer is testing: Alert prioritization, paging discipline, incident response.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Severity should map to required human action.

  • Page immediately: User-facing outage, fast error budget burn, data loss risk, security-impacting production issue.
  • Urgent ticket or chat: Degradation that needs same-day action but is not actively hurting users.
  • Backlog ticket: Capacity trend, cleanup task, non-production issue, or informational warning.
2️⃣

Remediation & Permanent Safeguards

I would define levels like: Each alert should include owner, service, impact, runbook, dashboard link, and escalation path. If no one needs to wake up and act immediately, it should not be a paging alert. This reduces alert fatigue and makes critical pages meaningful again.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Page immediately: User-facing outage, fast error budget burn, data loss risk, security-impacting production issue.."
⚡ 60-Second Elevator Pitch Talking Points
  • Page immediately: User-facing outage, fast error budget burn, data loss risk, security-impacting ...
  • Urgent ticket or chat: Degradation that needs same-day action but is not actively hurting users.
  • Backlog ticket: Capacity trend, cleanup task, non-production issue, or informational warning.
Advertisement
Want more Observability scenarios?
Explore our complete collection of scenario-based Observability interview runbooks.
Browse All Observability Questions →

📚 Related Production Scenarios in Observability