Q: Process is consuming high CPU — which commands would you use?
Full command toolkit for diagnosing high CPU processes: thread-level drilldown (top -H), syscall tracing (strace), kernel profiling (perf), and graceful mitigation.
#Linux #top #ps #strace #perf #jstack #CPU
🎙️ Candidate Opening & Architectural Context
"When a process consumes high CPU, my goal is to drill down from the system process level down to the exact thread, system call, or application function causing the burn."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Identify Process & System Context
Find the PID and understand its resource footprint:
top -c: PressPto sort by CPU. Shows full command-line paths.ps aux --sort=-%cpu | head -10: Fast scriptable capture of top 10 CPU consumers.pidstat 1 5 -u: Granular per-process CPU statistics over 5 seconds.
2️⃣
Thread-Level Inspection (top -H)
Multi-threaded runtimes (Java, Go, C++, Node worker threads) distribute load across threads:
top -H -p <PID>: Crucial command — displays individual threads inside the process as if they were processes. Identify the specific Thread ID (TID) pinned at 100%.ps -T -p <PID>: List all threads and their CPU consumption.- For Java: Convert TID to hexadecimal (
printf '%x\n' <TID>) and grep inside ajstack <PID>thread dump to find the exact line of Java code that is looping!
3️⃣
Trace Syscalls & Profile CPU (strace / perf)
Determine if CPU is burned in user code or kernel syscalls:
strace -p <PID> -c: Run for 10 seconds. Summarizes system calls, error rates, and time spent. If 90% time is infutex, process is suffering thread lock contention.strace -p <PID> -t -e trace=all: View live system calls in real-time.- Kernel Profiling (
perf): Runperf top -p <PID>orperf record -F 99 -p <PID> -g -- sleep 10 && perf reportto generate a call graph / flamegraph showing exact function hotspots.
4️⃣
Mitigate and Control Priority
Managing the process safely:
- Lower process priority:
renice +10 -p <PID>gives other critical system processes CPU priority. - Graceful shutdown:
kill -15 <PID>(SIGTERM). Only usekill -9as a last resort. - Restrict with cgroups: Apply CPU bandwidth limits via systemd slice or Docker
--cpus=1.5.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Use 'top -c' to find PID, 'top -H -p <PID>' to find the specific thread, and 'strace -c' or 'perf top' to pinpoint the looping function or lock contention."
⚡ 60-Second Elevator Pitch Talking Points
- Find process: 'top -c' (sort by P) or 'ps aux --sort=-%cpu | head -10' or 'pidstat 1 5 -u'.
- Inspect threads: 'top -H -p <PID>' to find the exact thread ID (TID) burning CPU.
- Trace syscalls: 'strace -p <PID> -c' to see time spent in kernel syscalls vs user code.
- Profile CPU: 'perf top -p <PID>' or capture thread dumps (jstack/pprof) before killing.
- Control: 'renice +10 <PID>' to deprioritize, or 'kill -15 <PID>' for graceful termination.
- Prevent: Set cgroup / systemd CPUQuota limits and container resource limits.
Advertisement