⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Senior DevOps / SRE Linux Performance & Processes Core Linux Skills

Q: Process is consuming high CPU — which commands would you use?

Full command toolkit for diagnosing high CPU processes: thread-level drilldown (top -H), syscall tracing (strace), kernel profiling (perf), and graceful mitigation.

#Linux #top #ps #strace #perf #jstack #CPU
🎙️ Candidate Opening & Architectural Context
"When a process consumes high CPU, my goal is to drill down from the system process level down to the exact thread, system call, or application function causing the burn."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Identify Process & System Context

Find the PID and understand its resource footprint:

  • top -c: Press P to sort by CPU. Shows full command-line paths.
  • ps aux --sort=-%cpu | head -10: Fast scriptable capture of top 10 CPU consumers.
  • pidstat 1 5 -u: Granular per-process CPU statistics over 5 seconds.
2️⃣

Thread-Level Inspection (top -H)

Multi-threaded runtimes (Java, Go, C++, Node worker threads) distribute load across threads:

  • top -H -p <PID>: Crucial command — displays individual threads inside the process as if they were processes. Identify the specific Thread ID (TID) pinned at 100%.
  • ps -T -p <PID>: List all threads and their CPU consumption.
  • For Java: Convert TID to hexadecimal (printf '%x\n' <TID>) and grep inside a jstack <PID> thread dump to find the exact line of Java code that is looping!
3️⃣

Trace Syscalls & Profile CPU (strace / perf)

Determine if CPU is burned in user code or kernel syscalls:

  • strace -p <PID> -c: Run for 10 seconds. Summarizes system calls, error rates, and time spent. If 90% time is in futex, process is suffering thread lock contention.
  • strace -p <PID> -t -e trace=all: View live system calls in real-time.
  • Kernel Profiling (perf): Run perf top -p <PID> or perf record -F 99 -p <PID> -g -- sleep 10 && perf report to generate a call graph / flamegraph showing exact function hotspots.
4️⃣

Mitigate and Control Priority

Managing the process safely:

  • Lower process priority: renice +10 -p <PID> gives other critical system processes CPU priority.
  • Graceful shutdown: kill -15 <PID> (SIGTERM). Only use kill -9 as a last resort.
  • Restrict with cgroups: Apply CPU bandwidth limits via systemd slice or Docker --cpus=1.5.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Use 'top -c' to find PID, 'top -H -p <PID>' to find the specific thread, and 'strace -c' or 'perf top' to pinpoint the looping function or lock contention."
⚡ 60-Second Elevator Pitch Talking Points
  • Find process: 'top -c' (sort by P) or 'ps aux --sort=-%cpu | head -10' or 'pidstat 1 5 -u'.
  • Inspect threads: 'top -H -p <PID>' to find the exact thread ID (TID) burning CPU.
  • Trace syscalls: 'strace -p <PID> -c' to see time spent in kernel syscalls vs user code.
  • Profile CPU: 'perf top -p <PID>' or capture thread dumps (jstack/pprof) before killing.
  • Control: 'renice +10 <PID>' to deprioritize, or 'kill -15 <PID>' for graceful termination.
  • Prevent: Set cgroup / systemd CPUQuota limits and container resource limits.
Advertisement
Want more Linux scenarios?
Explore our complete collection of scenario-based Linux interview runbooks.
Browse All Linux Questions →

📚 Related Production Scenarios in Linux