Q: What is the difference between CPU throttling and OOM Killed?
Deep architectural comparison between Kubernetes CPU Throttling and OOMKilled evictions, explaining how the Linux kernel enforces CPU CFS bandwidth versus memory hard limits.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
CPU Throttling Mechanics (Linux CFS Quota)
- **Resource Type**: Compressible. - **Kernel Mechanism**: Enforced via Linux Completely Fair Scheduler (CFS) bandwidth control (`cpu.cfs_quota_us` and `cpu.cfs_period_us`, typically 100ms periods). - **Behavior**: If a container with a limit of `500m` uses 50ms of CPU time in the first 20ms of a 100ms period, the kernel puts the threads to sleep for the remaining 80ms. - **Symptoms**: Process does NOT crash or restart; pod stays `1/1 Running`, but request latency spikes severely (p99 degradation).
# PromQL to detect CPU Throttling percentage:
sum(rate(container_cpu_cfs_throttled_periods_total[5m])) by (pod)
/
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod) * 100
OOMKilled Mechanics (Linux Cgroups OOM Killer)
- **Resource Type**: Non-compressible. - **Kernel Mechanism**: Enforced via `memory.max` in cgroup v2 (or `memory.limit_in_bytes` in cgroup v1). - **Behavior**: When container processes allocate physical RAM beyond `resources.limits.memory`, the kernel cannot 'throttle' memory. The kernel OOM killer sends `SIGKILL` (exit code 137). - **Symptoms**: Container crashes immediately, pod enters `CrashLoopBackOff`, and `kubectl describe pod` displays `OOMKilled: true` (Exit Code 137).
kubectl describe pod <pod-name>
# Look at Last State:
# Terminated: OOMKilled (Exit Code 137)
Direct Comparison Matrix
- **Resource**: CPU (Compressible) vs. Memory (Non-Compressible). - **Action on Breach**: Execution paused/slowed vs. Process killed violently (`SIGKILL`). - **Pod State**: Pod remains `Running` vs. Pod restarts with `CrashLoopBackOff`. - **Exit Code**: N/A (no exit) vs. `Exit Code 137`. - **Impact**: Increased request latency vs. Service outage / dropped connections.
Remediation Strategies
- **For CPU Throttling**: Increase CPU limits, remove CPU limits entirely (relying on requests), or optimize code concurrency. - **For OOMKilled**: Increase memory limits, profile heap memory leaks, tune JVM `-Xmx` to 75% of container memory limits, or fix memory bloat.
- CPU is compressible; exceeding CPU limits causes the kernel to throttle execution, slowing response times without killing the pod.
- Memory is non-compressible; exceeding memory limits triggers the Linux OOM killer, killing the container with Exit Code 137.
- Detect CPU throttling via container_cpu_cfs_throttled_periods_total in Prometheus.
- Detect OOMKilled via kubectl describe pod (Last State: OOMKilled).
- Consider removing CPU limits to eliminate latency jitter while always enforcing strict memory limits.