Q: A critical service suddenly starts consuming excessive CPU and memory. What would your troubleshooting methodology look like? How would you distinguish between application-level problems, infrastructure issues, resource misconfiguration, and traffic anomalies?
Deep diagnostic framework for distinguishing application memory leaks, infinite loops, CFS throttling, external database deadlocks, and traffic anomalies under pressure.
#Performance #Memory Leak #CPU Burn #Profiling #JVM #Go #strace #SRE
🎙️ Candidate Opening & Architectural Context
"When a service suddenly spikes in both CPU and memory, engineers often assume traffic surge or code bug. A rigorous SRE methodology classifies the anomaly across 4 distinct buckets before taking action."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Step 1: Classify the Anomaly Across 4 Buckets
Telemetry triage in the first 2 minutes:
- Bucket 1: Traffic Anomaly (Legit vs DDoS): Check Ingress RPS and active connections. If RPS jumped 10x proportionally with CPU, it's traffic load → Scale out horizontally (HPA / ASG).
- Bucket 2: Application-Level Bug: Traffic is flat, but CPU is pegged at 100% and memory climbs steadily without dropping → Infinite loop, regex catastrophic backtracking (ReDoS), or memory leak.
- Bucket 3: Downstream Bottleneck (Thread Pileup): Database or third-party payment gateway is slow/locked. Requests pile up waiting for I/O, exhausting thread pools and accumulating in-memory buffers.
- Bucket 4: Resource Misconfiguration: Low CPU limits causing Linux CFS throttling; or JVM max heap (
-Xmx) exceeding container memory limit, causing kernel OOMKill.
2️⃣
Step 2: CPU Deep-Dive (Thread & Syscall Profiling)
Isolate user code vs kernel syscalls:
top -H -p <PID>: Identify the specific Thread ID (TID) consuming CPU.strace -p <PID> -c: Check time spent in syscalls. If high infutex, service is suffering from thread lock contention.- Capture profiling artifacts before killing: Run
perf top, capture Javajstack <PID>, or profile Go runtime withpprof.
3️⃣
Step 3: Memory Deep-Dive (Leak vs Cache)
Distinguish healthy cache from dangerous memory leaks:
- Garbage Collection Metrics: Check GC frequency and duration in Prometheus. If Major/Full GC runs constantly and memory never recovers, it is an active Memory Leak (unbounded static collections, unclosed streams, or thread-local leaks).
- Capture Heap Dump:
jmap -dump:live,format=b,file=heap.bin <PID>or Gopprof/heap. - Analyze heap dump in Eclipse MAT (Memory Analyzer Tool) to find the largest memory-retaining objects.
4️⃣
Step 4: Mitigate Safely and Protect Customers
Immediate stabilization and safeguards:
- If thread pileup from downstream failure: Enable Circuit Breakers (Resilience4j / Envoy) and aggressive connection timeouts to fail fast.
- If memory leak: Temporarily increase memory limits or restart pods in a rolling fashion while developers analyze the heap dump.
- Tune cgroups: Remove artificial CPU limits that trigger kernel throttling if average node utilization is healthy.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Classify first: Flat traffic + pegged CPU = application bug/infinite loop. Flat traffic + rising memory without GC drop = memory leak. High response time + rising memory = downstream database bottleneck piling up threads."
⚡ 60-Second Elevator Pitch Talking Points
- Classify into 4 buckets: Traffic surge, application bug (leak/loop), downstream bottleneck (thread pileup), or misconfiguration.
- CPU drilldown: 'top -H -p <PID>' to find rogue thread; capture thread dump (jstack/pprof) or strace before killing.
- Memory drilldown: Monitor GC pauses in Prometheus. If full GC doesn't reclaim memory, capture heap dump (jmap/heap profile).
- Downstream check: Verify if slow database queries are causing worker thread accumulation and buffer exhaustion.
- Mitigate: Trip circuit breakers, scale out horizontally, or roll restart pods with captured diagnostic dumps.
Advertisement