⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Kubernetes Performance & Troubleshooting Performance Triage

Q: How would you troubleshoot high CPU or memory usage in Kubernetes?

End-to-end investigative procedure for identifying and resolving runaway CPU and memory saturation in Kubernetes clusters: isolating nodes vs pods, diagnosing memory leaks vs legitimate load, and application profiling.

#Kubernetes #CPU #Memory #kubectl top #Profiling #pprof #cgroups
🎙️ Candidate Opening & Architectural Context
"When a high CPU or memory alert triggers in Kubernetes, my workflow moves top-down: Cluster/Node level -> Pod level -> Container level -> Application runtime profiling (Go pprof / Java jstack)."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Step 1: Isolate Affected Nodes and Pods

Identify where the saturation is concentrated:

  • kubectl top nodes: Pinpoints if a single node is saturated or if the entire cluster is running out of headroom.
  • kubectl top pods -A --sort-by=cpu: Instantly lists the top CPU-consuming pods across all namespaces.
  • kubectl top pods -A --sort-by=memory: Instantly lists the top memory-consuming pods.
  • Check if the pod is near its configured limit: compare kubectl top pod <pod> with limits.cpu / limits.memory in kubectl describe pod.
2️⃣

Step 2: Differentiate CPU Spikes vs Memory Leaks

Analyze Grafana/Prometheus metrics to identify the pattern:

  • CPU Saturation: Correlate with request rate (RPS) in ingress metrics. If RPS spiked 5x, it's legitimate traffic -> trigger HPA or scale replicas. If RPS is flat but CPU hit 100%, it's an infinite loop, thread lock, or regex catastrophic backtracking.
  • Memory Saturation: Look at the memory graph over 24-48 hours. If memory steadily climbs in a sawtooth pattern and never drops after garbage collection, it is a Memory Leak.
3️⃣

Step 3: Capture Runtime Telemetry (Don't Restart Blindly)

Capture the root cause before restarting the container:

  • Java / JVM: Exec into pod or use ephemeral container: jcmd <pid> GC.heap_dump /tmp/dump.hprof and jstack <pid> to see locked threads.
  • Golang / Node.js: Query the pprof endpoint: curl http://localhost:6060/debug/pprof/profile?seconds=30 or heap profile.
  • Immediate Mitigation: Once profiling is captured, scale replicas horizontally (HPA) or do a rolling restart (kubectl rollout restart deployment/<app>) to restore customer SLA.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Diagnose top-down: 'kubectl top nodes' -> 'kubectl top pods --sort-by=cpu/memory' -> Correlate with traffic RPS in Grafana. For CPU without traffic spikes, capture thread dumps; for climbing memory, capture heap dumps before restarting."
⚡ 60-Second Elevator Pitch Talking Points
  • Identify culprit: 'kubectl top nodes' followed by 'kubectl top pods -A --sort-by=cpu / --sort-by=memory'.
  • Correlate: Compare with Grafana request rate (RPS). Traffic spike = scale replicas. Flat traffic + 100% CPU = code deadlock/infinite loop.
  • Memory triage: Check memory graph slope. Gradual steady climb without GC recovery = Memory Leak.
  • Capture telemetry before restart: Run jstack/jcmd for Java, or pprof heap profiles for Go.
  • Mitigate: Trigger HPA scale-out, or perform 'kubectl rollout restart' to buy time while developers fix the leak.
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes