⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Docker & Containers Interview Questions Scenario 110 of 158 in Docker & Containers
Staff Infrastructure Architect Docker Container Runtime & Systems Engineering Production Scenario

Q: Your microservices running in Docker and containerd frequently get killed by the Linux kernel Out-Of-Memory (OOM) killer during traffic bursts. In cgroups v1, container memory limits were binary (`memory.limit_in_bytes`), causing instant SIGKILL (exit code 137) as soon as usage exceeded the limit by a single byte. With your Linux hosts upgraded to kernel 6.x and cgroups v2, you want to implement soft memory limits and proactive throttling (`memory.high`) to allow applications to shed caches and throttle gracefully before hitting hard `memory.max` OOM termination.

Configure and tune Linux cgroups v2 memory controllers for containers. Understand the operational mechanics of `memory.high` proactive throttling versus `memory.max` hard OOM kills.

#Docker #Linux #Performance #cgroups #containerd
🎙️ Candidate Opening & Architectural Context
"Configure and tune Linux cgroups v2 memory controllers for containers. Understand the operational mechanics of `memory.high` proactive throttling versus `memory.max` hard OOM kills."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Docker Certified Associate (DCA) Hands-On Lab Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

Step 1

Verify Host cgroups v2 Unified Hierarchy

Verify that your Linux distribution has mounted the unified cgroups v2 hierarchy (`/sys/fs/cgroup`) rather than legacy cgroups v1 hybrid mounts.

# Check mounted cgroup filesystem type
stat -fc %T /sys/fs/cgroup/
# Expected output for cgroups v2: cgroup2fs

# Verify containerd/docker cgroup driver
docker info | grep "Cgroup Version"
# Output: Cgroup Version: 2
Pro Tip: Verify Host cgroups v2 Unified Hierarchy
Step 2

Compare cgroups v1 vs cgroups v2 Memory Knobs

Understand the architectural shifts: cgroups v1 mixed page cache and anonymous memory into a single rigid ceiling (`memory.limit_in_bytes`). cgroups v2 separates memory pressure levels into `memory.min` (hard protection), `memory.low` (soft protection), `memory.high` (proactive throttling), and `memory.max` (hard ceiling causing OOM kill).

<!-- cgroups v2 Memory Hierarchy -->
Host / Container Memory Space
  ├── memory.max  (Hard ceiling: Exceeding triggers kernel OOM killer -> Exit 137)
  ├── memory.high (Throttling zone: Processes throttled & forced into direct reclaim)
  ├── memory.low  (Best-effort protection: Reclaimed only under extreme host pressure)
  └── memory.min  (Hard guarantee: Never reclaimed by kernel page scanner)
Pro Tip: Compare cgroups v1 vs cgroups v2 Memory Knobs
Advertisement
Step 3

Configure Docker and containerd with memory.high Throttling

Launch containers specifying both soft throttling thresholds and hard maximum ceilings. When memory crosses `--memory-reservation` (which maps to `memory.high` in v2), the kernel slows down the container's processes and triggers aggressive reclaim rather than terminating the container.

# Docker run with memory reservation (memory.high) and hard limit (memory.max)
docker run -d \
  --name payment-service \
  --memory="2g" \
  --memory-reservation="1.6g" \
  --memory-swap="2g" \
  payment-api:v2.1

# Inspect container cgroup v2 controller files
CONTAINER_ID=$(docker inspect payment-service --format '{{.Id}}')
cat /sys/fs/cgroup/docker/${CONTAINER_ID}/memory.high
cat /sys/fs/cgroup/docker/${CONTAINER_ID}/memory.max
cat /sys/fs/cgroup/docker/${CONTAINER_ID}/memory.events
Pro Tip: Configure Docker and containerd with memory.high Throttling
Step 4

Monitor Memory Pressure Events and Reclaim Stall Time

Read `/sys/fs/cgroup/.../memory.events` to track `high`, `max`, and `oom` event counters. Inspect `memory.pressure` (PSI - Pressure Stall Information) to measure the percentage of time processes stall waiting for memory pages.

# Check container memory events
cat /sys/fs/cgroup/docker/${CONTAINER_ID}/memory.events
# Output:
# low 0
# high 42
# max 0
# oom 0
# oom_kill 0

# Inspect Pressure Stall Information (PSI)
cat /sys/fs/cgroup/docker/${CONTAINER_ID}/memory.pressure
Pro Tip: Monitor Memory Pressure Events and Reclaim Stall Time
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"cgroups v2 introduces `memory.high` as a throttling and proactive reclaim boundary. Setting `memory.high` below `memory.max` gives containers breathing room to flush caches and throttle throughput under memory spikes instead of suffering catastrophic OOM kills."
⚡ 60-Second Elevator Pitch Talking Points
  • U
  • p
  • g
  • r
  • a
  • d
  • i
  • n
  • g
  • t
  • o
  • c
  • g
  • r
  • o
  • u
  • p
  • s
  • v
  • 2
  • e
  • n
  • a
  • b
  • l
  • e
  • d
  • u
  • s
  • t
  • o
  • e
  • l
  • i
  • m
  • i
  • n
  • a
  • t
  • e
  • s
  • u
  • d
  • d
  • e
  • n
  • O
  • O
  • M
  • k
  • i
  • l
  • l
  • s
  • a
  • c
  • r
  • o
  • s
  • s
  • o
  • u
  • r
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • f
  • l
  • e
  • e
  • t
  • .
  • B
  • y
  • c
  • o
  • n
  • f
  • i
  • g
  • u
  • r
  • i
  • n
  • g
  • `
  • m
  • e
  • m
  • o
  • r
  • y
  • .
  • h
  • i
  • g
  • h
  • `
  • a
  • l
  • o
  • n
  • g
  • s
  • i
  • d
  • e
  • `
  • m
  • e
  • m
  • o
  • r
  • y
  • .
  • m
  • a
  • x
  • `
  • ,
  • t
  • h
  • e
  • L
  • i
  • n
  • u
  • x
  • k
  • e
  • r
  • n
  • e
  • l
  • a
  • p
  • p
  • l
  • i
  • e
  • s
  • p
  • r
  • o
  • g
  • r
  • e
  • s
  • s
  • i
  • v
  • e
  • t
  • h
  • r
  • o
  • t
  • t
  • l
  • i
  • n
  • g
  • a
  • n
  • d
  • d
  • i
  • r
  • e
  • c
  • t
  • p
  • a
  • g
  • e
  • r
  • e
  • c
  • l
  • a
  • i
  • m
  • w
  • h
  • e
  • n
  • m
  • e
  • m
  • o
  • r
  • y
  • s
  • p
  • i
  • k
  • e
  • s
  • o
  • c
  • c
  • u
  • r
  • ,
  • a
  • l
  • l
  • o
  • w
  • i
  • n
  • g
  • s
  • e
  • r
  • v
  • i
  • c
  • e
  • s
  • t
  • o
  • s
  • h
  • e
  • d
  • c
  • a
  • c
  • h
  • e
  • s
  • g
  • r
  • a
  • c
  • e
  • f
  • u
  • l
  • l
  • y
  • w
  • h
  • i
  • l
  • e
  • a
  • l
  • e
  • r
  • t
  • i
  • n
  • g
  • o
  • u
  • r
  • a
  • u
  • t
  • o
  • s
  • c
  • a
  • l
  • e
  • r
  • s
  • b
  • e
  • f
  • o
  • r
  • e
  • a
  • n
  • O
  • O
  • M
  • k
  • i
  • l
  • l
  • c
  • a
  • n
  • h
  • a
  • p
  • p
  • e
  • n
  • .
Advertisement
Want more Docker & Containers scenarios?
Explore our complete collection of scenario-based Docker & Containers interview runbooks.
Browse All Docker & Containers Questions →