Q: How do you configure CPU and memory requests/limits?
Definitive architectural guide for configuring CPU and memory requests/limits: how kube-scheduler uses requests, how the Linux kernel enforces limits, CFS CPU throttling vs OOMKilled (exit 137), and QoS classes.
#Kubernetes #Requests #Limits #QoS #OOMKilled #Throttling #cgroups
🎙️ Candidate Opening & Architectural Context
"Configuring requests and limits is not guess-work. Requests determine where the scheduler places pods; limits determine when the Linux kernel throttles or terminates them. Setting them incorrectly causes either cluster starvation or silent application slowness."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Requests vs Limits Mechanics
Understanding how the control plane and Linux kernel treat them differently:
⏱️ CPU (Compressible)
When a pod hits its CPU limit, it is NOT killed. The kernel CFS quota throttles CPU cycles, causing latency spikes and slow response times.
💥 Memory (Incompressible)
Memory cannot be throttled. When a container exceeds its memory limit, the Linux kernel OOM Killer terminates it immediately (Exit Code 137, OOMKilled).
- spec.resources.requests: The guaranteed reservation. The
kube-schedulercalculates node placement strictly based on sum of requests vs node capacity. If a node has 4 cores and 3.8 cores are requested, no new pod requesting 500m can fit. - spec.resources.limits: The hard ceiling. Enforced by Linux kernel cgroups. How the kernel responds depends on the resource type:
2️⃣
The 3 Kubernetes Quality of Service (QoS) Classes
Kubernetes automatically assigns QoS based on requests and limits:
- 1. Guaranteed (Highest Priority):
requests.cpu == limits.cpuANDrequests.memory == limits.memoryfor all containers. Last to be evicted when node is under resource pressure. Best for critical databases and core services. - 2. Burstable (Medium Priority): Requests are less than limits (e.g. request 500m, limit 2000m). Allowed to burst when extra node capacity exists. Standard for most web APIs.
- 3. BestEffort (Lowest Priority): No requests and no limits set. First to be killed instantly when a node experiences memory pressure.
3️⃣
Production Right-Sizing Best Practices
How senior SREs avoid common traps:
- Right-Sizing: Use Prometheus historical p95 usage or Vertical Pod Autoscaler (VPA in recommendation mode) to base requests on actual p95 peak load + 20% headroom.
- The CPU Limit Debate: Many high-scale teams (Google, Uber) remove CPU limits on latency-sensitive Burstable services to prevent CFS throttle penalties during micro-bursts, relying on HPA to scale out.
- Memory: Always set memory limits equal or close to requests to prevent noisy neighbors from consuming all node RAM.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Requests are for scheduling (reservations); limits are for kernel enforcement. CPU is compressible (hits limit -> throttles); memory is incompressible (hits limit -> OOMKilled Exit 137). Set requests = limits on databases for Guaranteed QoS, and right-size requests to p95 peak usage."
⚡ 60-Second Elevator Pitch Talking Points
- Requests: Guaranteed minimum reserved by kube-scheduler; placement depends on it.
- Limits: Maximum ceiling enforced by cgroups.
- CPU behavior: Compressible -> throttled via CFS quota (app slows down, no crash).
- Memory behavior: Incompressible -> killed immediately by OOM Killer (Exit Code 137).
- QoS Classes: Guaranteed (requests == limits, safest), Burstable (requests < limits), BestEffort (no values, evicted first).
- Best practice: Right-size with VPA recommendation mode; always set memory limits.
Advertisement