⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 184 of 194 in Kubernetes
Senior SRE / Systems Engineer Kubernetes Pod Performance & Profiling J.P. Morgan Technical Loop

Q: You see high CPU usage in one pod, but logs look clean. What next?

Root cause analysis workflow for diagnosing a Kubernetes pod pinning 100% CPU when application logs contain zero errors or anomalies.

#Kubernetes #Linux #CPU Throttling #Profiling #JVM #Async #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When a pod consumes 100% CPU while logging nothing, the application is stuck in a silent CPU-bound spin: an infinite loop, JVM garbage collection thrashing (Stop-The-World full GC), regex catastrophic backtracking, deadlock spinning on busy-wait locks, or an unhandled exception loop that silently swallows errors."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Identify Whether All Cores or Single Core Are Pinned

Exec into the pod or the underlying host node. Run `top` or `htop` and press `1` to see per-core CPU usage, and `H` to show thread-level CPU usage. Identify the exact thread ID (TID) consuming the CPU.

kubectl exec -it <pod-name> -- top -H
# Note the high CPU PID/TID: e.g. PID 42 using 99.8% CPU
2

Capture Thread Dumps or CPU Stack Traces

Convert the thread ID to hexadecimal and inspect thread stacks: - **Java / JVM**: Use `jstack ` and search for the hexadecimal TID (`nid=0x2a`). - **Go / Node / Python**: Capture a pprof or async-profiler flamegraph.

# For JVM:
printf "%x\n" 42  # Output: 2a
jstack <pid> | grep -A 20 "nid=0x2a"
# For Go:
curl http://localhost:6060/debug/pprof/profile?seconds=30 > cpu.pprof
go tool pprof -http=:8080 cpu.pprof
Advertisement
3

Check Garbage Collection Thrashing & Memory Pressure

Inspect memory allocation. If the application heap is 98% full and JVM Garbage Collector threads (`VM Thread`, `GC task thread#0`) are running continuously attempting to reclaim unreachable objects, CPU will pin at 100% while application code is suspended.

jstat -gcutil <pid> 1000 5
# Inspect FGC (Full GC count) and FGCT (Full GC time)
4

Inspect Kernel Syscalls with strace / perf

If thread dumps cannot be gathered, attach `perf` or `strace` from the node to see what system calls the thread is executing (e.g. spinning on `futex` or non-blocking socket reads).

Pro Tip: Production Tip: Always capture thread dumps and profiler flamegraphs BEFORE killing or restarting the pod, otherwise the root cause evidence is permanently lost.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Clean logs with high CPU point to silent CPU loops, regex backtracking, or GC thrashing. Capture thread dumps (jstack) or CPU profilers (pprof/perf) to identify the spinning thread ID before restarting."
⚡ 60-Second Elevator Pitch Talking Points
  • Run top -H inside the container to identify the exact thread ID (TID) consuming CPU.
  • Capture thread dumps (jstack) and correlate the hex TID to find the exact code line.
  • Verify memory utilization to rule out continuous JVM Garbage Collection thrashing.
  • Preserve thread dumps and heap telemetry for post-incident root cause analysis before pod restarts.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →