⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff / Principal SRE SRE & Operations Performance Engineering Incident Triage

Q: A critical service suddenly starts consuming excessive CPU and memory. What would your troubleshooting methodology look like? How would you distinguish between application-level problems, infrastructure issues, resource misconfiguration, and traffic anomalies?

Deep diagnostic framework for distinguishing application memory leaks, infinite loops, CFS throttling, external database deadlocks, and traffic anomalies under pressure.

#Performance #Memory Leak #CPU Burn #Profiling #JVM #Go #strace #SRE
🎙️ Candidate Opening & Architectural Context
"When a service suddenly spikes in both CPU and memory, engineers often assume traffic surge or code bug. A rigorous SRE methodology classifies the anomaly across 4 distinct buckets before taking action."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Step 1: Classify the Anomaly Across 4 Buckets

Telemetry triage in the first 2 minutes:

  • Bucket 1: Traffic Anomaly (Legit vs DDoS): Check Ingress RPS and active connections. If RPS jumped 10x proportionally with CPU, it's traffic load → Scale out horizontally (HPA / ASG).
  • Bucket 2: Application-Level Bug: Traffic is flat, but CPU is pegged at 100% and memory climbs steadily without dropping → Infinite loop, regex catastrophic backtracking (ReDoS), or memory leak.
  • Bucket 3: Downstream Bottleneck (Thread Pileup): Database or third-party payment gateway is slow/locked. Requests pile up waiting for I/O, exhausting thread pools and accumulating in-memory buffers.
  • Bucket 4: Resource Misconfiguration: Low CPU limits causing Linux CFS throttling; or JVM max heap (-Xmx) exceeding container memory limit, causing kernel OOMKill.
2️⃣

Step 2: CPU Deep-Dive (Thread & Syscall Profiling)

Isolate user code vs kernel syscalls:

  • top -H -p <PID>: Identify the specific Thread ID (TID) consuming CPU.
  • strace -p <PID> -c: Check time spent in syscalls. If high in futex, service is suffering from thread lock contention.
  • Capture profiling artifacts before killing: Run perf top, capture Java jstack <PID>, or profile Go runtime with pprof.
3️⃣

Step 3: Memory Deep-Dive (Leak vs Cache)

Distinguish healthy cache from dangerous memory leaks:

  • Garbage Collection Metrics: Check GC frequency and duration in Prometheus. If Major/Full GC runs constantly and memory never recovers, it is an active Memory Leak (unbounded static collections, unclosed streams, or thread-local leaks).
  • Capture Heap Dump: jmap -dump:live,format=b,file=heap.bin <PID> or Go pprof/heap.
  • Analyze heap dump in Eclipse MAT (Memory Analyzer Tool) to find the largest memory-retaining objects.
4️⃣

Step 4: Mitigate Safely and Protect Customers

Immediate stabilization and safeguards:

  • If thread pileup from downstream failure: Enable Circuit Breakers (Resilience4j / Envoy) and aggressive connection timeouts to fail fast.
  • If memory leak: Temporarily increase memory limits or restart pods in a rolling fashion while developers analyze the heap dump.
  • Tune cgroups: Remove artificial CPU limits that trigger kernel throttling if average node utilization is healthy.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Classify first: Flat traffic + pegged CPU = application bug/infinite loop. Flat traffic + rising memory without GC drop = memory leak. High response time + rising memory = downstream database bottleneck piling up threads."
⚡ 60-Second Elevator Pitch Talking Points
  • Classify into 4 buckets: Traffic surge, application bug (leak/loop), downstream bottleneck (thread pileup), or misconfiguration.
  • CPU drilldown: 'top -H -p <PID>' to find rogue thread; capture thread dump (jstack/pprof) or strace before killing.
  • Memory drilldown: Monitor GC pauses in Prometheus. If full GC doesn't reclaim memory, capture heap dump (jmap/heap profile).
  • Downstream check: Verify if slow database queries are causing worker thread accumulation and buffer exhaustion.
  • Mitigate: Trip circuit breakers, scale out horizontally, or roll restart pods with captured diagnostic dumps.
Advertisement
Want more SRE & Operations scenarios?
Explore our complete collection of scenario-based SRE & Operations interview runbooks.
Browse All SRE & Operations Questions →

📚 Related Production Scenarios in SRE & Operations