⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Docker & Containers Interview Questions Scenario 136 of 158 in Docker & Containers
Staff Infrastructure Architect Docker Container Runtime & Systems Engineering Production Scenario

Q: Your monitoring alerts trigger high file descriptor (FD) and process table exhaustion on several Kubernetes worker nodes. Running `ps aux | grep shim` reveals over 1,400 running `containerd-shim-runc-v2` processes, even though the node only hosts 40 active pods. Kubelet cannot schedule new containers. You must analyze why containerd creates a dedicated shim per container, triage the orphaned shims and leaked UNIX domain sockets, and restore the node to a clean operating state.

Deconstruct the containerd `shim-v2` architecture. Understand why a shim exists per container/pod sandbox, how it decouples runc execution, and how to triage orphaned shims and file descriptor leaks.

#Docker #containerd #Architecture #Linux #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"Deconstruct the containerd `shim-v2` architecture. Understand why a shim exists per container/pod sandbox, how it decouples runc execution, and how to triage orphaned shims and file descriptor leaks."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Docker Certified Associate (DCA) Hands-On Lab Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

Step 1

Deconstruct the Architectural Purpose of containerd-shim-v2

In OCI container runtimes, `runc` executes the container setup and exits. Without a shim, the daemon would have to remain the parent process of every container, preventing daemon restarts. The `shim-v2` stays alive to: act as sub-reaper for container processes, hold container PTY and stdio pipes open, report exit status codes, and allow the daemon to restart without terminating workloads.

<!-- containerd Architecture Flow -->
kubelet / docker CLI
  └── containerd (Daemon)
        └── containerd-shim-runc-v2 (1 per Pod/Container)
              ├── Holds stdio FIFOs & Unix Sockets
              ├── Monitors exit code
              └── runc (Runs setup, launches App, then exits!)
                    └── Application Process (Child of Shim)
Pro Tip: Deconstruct the Architectural Purpose of containerd-shim-v2
Step 2

Inspect Shim Communication via TTRPC and Unix Sockets

containerd communicates with shims via TTRPC (a lightweight gRPC alternative optimized for low-memory environments) over local UNIX domain sockets located in `/run/containerd/io.containerd.runtime.v2.task/`.

# Inspect active shim sockets
ls -l /run/containerd/io.containerd.runtime.v2.task/k8s.io/

# Inspect open file descriptors of a shim process
lsof -p $(pgrep -f containerd-shim-runc-v2 | head -n 1)
Pro Tip: Inspect Shim Communication via TTRPC and Unix Sockets
Advertisement
Step 3

Identify and Terminate Orphaned Shims

When a container process crashes violently or experiences an unhandled kernel fault, the shim may fail to clean up its state, remaining stuck in the process table. Identify shims whose child processes no longer exist.

# Find shims without any child processes
for shim_pid in $(pgrep -f containerd-shim-runc-v2); do
    CHILDREN=$(pgrep -P $shim_pid | wc -l)
    if [ "$CHILDREN" -eq 0 ]; then
        echo "Orphaned shim detected: PID $shim_pid"
        # Inspect shim working directory
        ls -l /proc/$shim_pid/cwd
    fi
done

# Safely kill orphaned shims
# kill -9 <ORPHAN_PID>
Pro Tip: Identify and Terminate Orphaned Shims
Step 4

Clean Leaked Runtime Sockets and State Bundles

Remove stale task directories from `/run/containerd` and purge orphaned FIFOs using `ctr tasks delete` to reclaim leaked file descriptors.

# Delete dead tasks directly in containerd
ctr -n k8s.io tasks delete -f <TASK_ID>

# Verify shim count returns to matching active pod count
pgrep -c -f containerd-shim-runc-v2
Pro Tip: Clean Leaked Runtime Sockets and State Bundles
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"`containerd-shim-v2` decouples container lifecycles from the main containerd daemon using TTRPC and dedicated UNIX sockets. Orphaned shims occur when child processes exit without signaling the runtime, causing process table and file descriptor leaks that must be purged via runtime task cleanup."
⚡ 60-Second Elevator Pitch Talking Points
  • W
  • e
  • d
  • i
  • a
  • g
  • n
  • o
  • s
  • e
  • d
  • a
  • s
  • e
  • v
  • e
  • r
  • e
  • f
  • i
  • l
  • e
  • d
  • e
  • s
  • c
  • r
  • i
  • p
  • t
  • o
  • r
  • e
  • x
  • h
  • a
  • u
  • s
  • t
  • i
  • o
  • n
  • i
  • s
  • s
  • u
  • e
  • b
  • y
  • t
  • r
  • a
  • c
  • i
  • n
  • g
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • d
  • r
  • u
  • n
  • t
  • i
  • m
  • e
  • i
  • n
  • t
  • e
  • r
  • n
  • a
  • l
  • s
  • .
  • O
  • v
  • e
  • r
  • 1
  • ,
  • 0
  • 0
  • 0
  • o
  • r
  • p
  • h
  • a
  • n
  • e
  • d
  • `
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • d
  • -
  • s
  • h
  • i
  • m
  • -
  • r
  • u
  • n
  • c
  • -
  • v
  • 2
  • `
  • p
  • r
  • o
  • c
  • e
  • s
  • s
  • e
  • s
  • h
  • a
  • d
  • l
  • e
  • a
  • k
  • e
  • d
  • d
  • u
  • e
  • t
  • o
  • u
  • n
  • h
  • a
  • n
  • d
  • l
  • e
  • d
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • c
  • r
  • a
  • s
  • h
  • e
  • s
  • .
  • W
  • e
  • s
  • c
  • r
  • i
  • p
  • t
  • e
  • d
  • a
  • n
  • a
  • u
  • t
  • o
  • m
  • a
  • t
  • e
  • d
  • d
  • e
  • t
  • e
  • c
  • t
  • i
  • o
  • n
  • l
  • o
  • o
  • p
  • t
  • h
  • a
  • t
  • i
  • d
  • e
  • n
  • t
  • i
  • f
  • i
  • e
  • s
  • s
  • h
  • i
  • m
  • s
  • w
  • i
  • t
  • h
  • z
  • e
  • r
  • o
  • c
  • h
  • i
  • l
  • d
  • p
  • r
  • o
  • c
  • e
  • s
  • s
  • e
  • s
  • ,
  • p
  • u
  • r
  • g
  • e
  • s
  • t
  • h
  • e
  • i
  • r
  • s
  • t
  • a
  • l
  • e
  • T
  • T
  • R
  • P
  • C
  • s
  • o
  • c
  • k
  • e
  • t
  • s
  • w
  • i
  • t
  • h
  • `
  • c
  • t
  • r
  • t
  • a
  • s
  • k
  • s
  • d
  • e
  • l
  • e
  • t
  • e
  • `
  • ,
  • a
  • n
  • d
  • b
  • r
  • o
  • u
  • g
  • h
  • t
  • p
  • r
  • o
  • c
  • e
  • s
  • s
  • c
  • o
  • u
  • n
  • t
  • s
  • b
  • a
  • c
  • k
  • i
  • n
  • t
  • o
  • a
  • l
  • i
  • g
  • n
  • m
  • e
  • n
  • t
  • .
Advertisement
Want more Docker & Containers scenarios?
Explore our complete collection of scenario-based Docker & Containers interview runbooks.
Browse All Docker & Containers Questions →