Q: A service uses 2GB of RAM. Do you alert when it hits 1.5GB (75%) or 1.9GB (95%)? Explain your reasoning.
A static threshold often fails because it ignores the rate of change.
#Observability #Observability #L1 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""During a high-traffic production event, our observability stack proved essential in isolating this latency surge. The interviewer is testing: Lead time, rate of change, threshold theory.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Production Solution & Architecture
A static threshold often fails because it ignores the *rate of change*. If memory is leaking slowly at 1 MB/hour, alerting at 75% gives me 500 hours to fix it — an annoying alert I don't need right now. If it spikes extremely fast, alerting at 95% might only give me 2 seconds before the OOM kill happens, making the alert useless because it's too late. The better approach is to alert on the Time To Exhaustion. I would use the Prometheus predict_linear() function over the last hour. If the slope indicates we will hit 100% in the next 4 hours, it alerts. This gives me actionable lead time, regardless of whether memory is currently at 40% or 90%.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: A static threshold often fails because it ignores the *rate of change*.."
⚡ 60-Second Elevator Pitch Talking Points
- Immediate Triage: A static threshold often fails because it ignores the rate of change.
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.
Advertisement