⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Kubernetes Interview Questions Scenario 194 of 194 in Kubernetes
Senior SRE / Linux Engineer Kubernetes Resource Management & Kernel Cgroups Operations & Support Loop

Q: What is the difference between CPU throttling and OOM Killed?

Deep architectural comparison between Kubernetes CPU Throttling and OOMKilled evictions, explaining how the Linux kernel enforces CPU CFS bandwidth versus memory hard limits.

#Kubernetes #Linux #cgroups #CPU Throttling #OOMKilled #CFS #Performance
🎙️ Candidate Opening & Architectural Context
"CPU and Memory are fundamentally different computing resources: CPU is compressible, while Memory is non-compressible. When a container exceeds its CPU limit, the Linux kernel throttles execution without terminating the process. When a container exceeds its Memory limit, the Linux kernel invokes the Out-of-Memory (OOM) killer and violently terminates the process (`SIGKILL`)."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

CPU Throttling Mechanics (Linux CFS Quota)

- **Resource Type**: Compressible. - **Kernel Mechanism**: Enforced via Linux Completely Fair Scheduler (CFS) bandwidth control (`cpu.cfs_quota_us` and `cpu.cfs_period_us`, typically 100ms periods). - **Behavior**: If a container with a limit of `500m` uses 50ms of CPU time in the first 20ms of a 100ms period, the kernel puts the threads to sleep for the remaining 80ms. - **Symptoms**: Process does NOT crash or restart; pod stays `1/1 Running`, but request latency spikes severely (p99 degradation).

# PromQL to detect CPU Throttling percentage:
sum(rate(container_cpu_cfs_throttled_periods_total[5m])) by (pod)
/
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod) * 100
2

OOMKilled Mechanics (Linux Cgroups OOM Killer)

- **Resource Type**: Non-compressible. - **Kernel Mechanism**: Enforced via `memory.max` in cgroup v2 (or `memory.limit_in_bytes` in cgroup v1). - **Behavior**: When container processes allocate physical RAM beyond `resources.limits.memory`, the kernel cannot 'throttle' memory. The kernel OOM killer sends `SIGKILL` (exit code 137). - **Symptoms**: Container crashes immediately, pod enters `CrashLoopBackOff`, and `kubectl describe pod` displays `OOMKilled: true` (Exit Code 137).

kubectl describe pod <pod-name>
# Look at Last State:
# Terminated: OOMKilled (Exit Code 137)
Advertisement
3

Direct Comparison Matrix

- **Resource**: CPU (Compressible) vs. Memory (Non-Compressible). - **Action on Breach**: Execution paused/slowed vs. Process killed violently (`SIGKILL`). - **Pod State**: Pod remains `Running` vs. Pod restarts with `CrashLoopBackOff`. - **Exit Code**: N/A (no exit) vs. `Exit Code 137`. - **Impact**: Increased request latency vs. Service outage / dropped connections.

Pro Tip: Industry Best Practice: Many senior SRE teams omit CPU limits entirely while setting generous CPU requests to prevent artificial CFS throttling, while ALWAYS strictly setting memory limits.
4

Remediation Strategies

- **For CPU Throttling**: Increase CPU limits, remove CPU limits entirely (relying on requests), or optimize code concurrency. - **For OOMKilled**: Increase memory limits, profile heap memory leaks, tune JVM `-Xmx` to 75% of container memory limits, or fix memory bloat.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"CPU is compressible: exceeding limits triggers CFS throttling (latency increases, pod does not die). Memory is non-compressible: exceeding limits triggers the kernel OOM killer (SIGKILL, exit code 137, pod restarts)."
⚡ 60-Second Elevator Pitch Talking Points
  • CPU is compressible; exceeding CPU limits causes the kernel to throttle execution, slowing response times without killing the pod.
  • Memory is non-compressible; exceeding memory limits triggers the Linux OOM killer, killing the container with Exit Code 137.
  • Detect CPU throttling via container_cpu_cfs_throttled_periods_total in Prometheus.
  • Detect OOMKilled via kubectl describe pod (Last State: OOMKilled).
  • Consider removing CPU limits to eliminate latency jitter while always enforcing strict memory limits.
Advertisement
📥 FREE DOWNLOAD · 101-PAGE COMPANION HANDBOOK
Studying for Kubernetes & SRE Technical Rounds?
Download the complete 100-question PDF field guide covering all 11 core modules with offline diagnostic runbooks.
📥 Download PDF (Free) Read Online Guide →
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →