Q: How would you troubleshoot high CPU or memory usage in Kubernetes?
End-to-end investigative procedure for identifying and resolving runaway CPU and memory saturation in Kubernetes clusters: isolating nodes vs pods, diagnosing memory leaks vs legitimate load, and application profiling.
#Kubernetes #CPU #Memory #kubectl top #Profiling #pprof #cgroups
🎙️ Candidate Opening & Architectural Context
"When a high CPU or memory alert triggers in Kubernetes, my workflow moves top-down: Cluster/Node level -> Pod level -> Container level -> Application runtime profiling (Go pprof / Java jstack)."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Step 1: Isolate Affected Nodes and Pods
Identify where the saturation is concentrated:
kubectl top nodes: Pinpoints if a single node is saturated or if the entire cluster is running out of headroom.kubectl top pods -A --sort-by=cpu: Instantly lists the top CPU-consuming pods across all namespaces.kubectl top pods -A --sort-by=memory: Instantly lists the top memory-consuming pods.- Check if the pod is near its configured limit: compare
kubectl top pod <pod>withlimits.cpu/limits.memoryinkubectl describe pod.
2️⃣
Step 2: Differentiate CPU Spikes vs Memory Leaks
Analyze Grafana/Prometheus metrics to identify the pattern:
- CPU Saturation: Correlate with request rate (RPS) in ingress metrics. If RPS spiked 5x, it's legitimate traffic -> trigger HPA or scale replicas. If RPS is flat but CPU hit 100%, it's an infinite loop, thread lock, or regex catastrophic backtracking.
- Memory Saturation: Look at the memory graph over 24-48 hours. If memory steadily climbs in a sawtooth pattern and never drops after garbage collection, it is a Memory Leak.
3️⃣
Step 3: Capture Runtime Telemetry (Don't Restart Blindly)
Capture the root cause before restarting the container:
- Java / JVM: Exec into pod or use ephemeral container:
jcmd <pid> GC.heap_dump /tmp/dump.hprofandjstack <pid>to see locked threads. - Golang / Node.js: Query the pprof endpoint:
curl http://localhost:6060/debug/pprof/profile?seconds=30or heap profile. - Immediate Mitigation: Once profiling is captured, scale replicas horizontally (HPA) or do a rolling restart (
kubectl rollout restart deployment/<app>) to restore customer SLA.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Diagnose top-down: 'kubectl top nodes' -> 'kubectl top pods --sort-by=cpu/memory' -> Correlate with traffic RPS in Grafana. For CPU without traffic spikes, capture thread dumps; for climbing memory, capture heap dumps before restarting."
⚡ 60-Second Elevator Pitch Talking Points
- Identify culprit: 'kubectl top nodes' followed by 'kubectl top pods -A --sort-by=cpu / --sort-by=memory'.
- Correlate: Compare with Grafana request rate (RPS). Traffic spike = scale replicas. Flat traffic + 100% CPU = code deadlock/infinite loop.
- Memory triage: Check memory graph slope. Gradual steady climb without GC recovery = Memory Leak.
- Capture telemetry before restart: Run jstack/jcmd for Java, or pprof heap profiles for Go.
- Mitigate: Trigger HPA scale-out, or perform 'kubectl rollout restart' to buy time while developers fix the leak.
Advertisement