Q: Your application latency has suddenly spiked but CPU, Memory, and Network I/O metrics remain normal. What do you check first?
If the application instance's local resources are fine, the latency is almost certainly caused by an external downstream dependency.
#Observability #Observability #L1 #Monitoring #Prometheus #SRE
🎙️ Candidate Opening & Architectural Context
""During a high-traffic production event, our observability stack proved essential in isolating this latency surge. The interviewer is testing: Understanding of external dependencies and basic troubleshooting workflow.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
If the application instance's local resources are fine, the latency is almost certainly caused by an external downstream dependency.
- Database query latency (e.g., table locks, missing indexes).
- Third-party API timeouts or throttling (e.g., Stripe, SendGrid).
- Cache latency (e.g., Redis cluster eviction policies taking too long or connection exhaustion).
2️⃣
Remediation & Permanent Safeguards
I would check APM (Application Performance Monitoring) distributed traces or dependency metrics. I am looking for: If external telemetry indicates they are fast, the application might be experiencing thread pool exhaustion or garbage collection (GC) pauses internally, which APM thread/GC metrics would reveal despite low host CPU.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Database query latency (e.g., table locks, missing indexes).."
⚡ 60-Second Elevator Pitch Talking Points
- Database query latency (e.g., table locks, missing indexes).
- Third-party API timeouts or throttling (e.g., Stripe, SendGrid).
- Cache latency (e.g., Redis cluster eviction policies taking too long or connection exhaustion).
Advertisement