⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Docker & Containers Interview Questions Scenario 109 of 158 in Docker & Containers
Staff Infrastructure Architect Docker Container Runtime & Systems Engineering Production Scenario

Q: A production Kubernetes node running containerd suddenly enters `NotReady` status. Kubelet logs report `rpc error: code = DeadlineExceeded desc = context deadline exceeded` while calling `RuntimeService.Status` on `/run/containerd/containerd.sock`. Running `crictl ps` hangs indefinitely, while host CPU and memory appear normal. You must systematically triage the containerd daemon, inspect shim processes, analyze gRPC socket health, and restore container runtime functionality without rebooting the bare-metal worker node.

Diagnose and resolve containerd CRI socket freezes (`/run/containerd/containerd.sock`), client timeouts in kubelet, and distinguish between low-level `ctr` vs Kubernetes-focused `crictl` tooling.

#Docker #containerd #Kubernetes #Linux #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"Diagnose and resolve containerd CRI socket freezes (`/run/containerd/containerd.sock`), client timeouts in kubelet, and distinguish between low-level `ctr` vs Kubernetes-focused `crictl` tooling."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Docker Certified Associate (DCA) Hands-On Lab Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

Step 1

Differentiate CLI Tools: crictl vs ctr vs nerdctl

Understand the tool boundaries: `crictl` communicates over the Kubernetes CRI gRPC socket (`/run/containerd/containerd.sock`) using the `k8s.io` namespace. `ctr` is a low-level debugging CLI shipped directly with containerd that bypasses the CRI plugin entirely. `nerdctl` provides a Docker-compatible CLI for containerd with compose and rootless support.

# crictl targets the Kubernetes CRI plugin
crictl --runtime-endpoint unix:///run/containerd/containerd.sock ps

# ctr operates directly on containerd internal namespaces (bypassing CRI)
ctr -n k8s.io containers list
ctr -n default images list

# Inspect containerd socket availability via curl/socat
echo -e "GET / HTTP/1.0\r\n" | socat - UNIX-CONNECT:/run/containerd/containerd.sock
Pro Tip: Differentiate CLI Tools: crictl vs ctr vs nerdctl
Step 2

Inspect containerd Daemon Thread Dumps and Tracing

Send `SIGUSR1` to the containerd process to trigger a stack trace dump into `journalctl` without terminating active workloads. Check for gRPC worker goroutines deadlocked on filesystem I/O locks or unresponsive shims.

# Send SIGUSR1 to generate goroutine stack trace
pkill -SIGUSR1 containerd

# Inspect stack traces in systemd journal
journalctl -u containerd -n 200 --no-pager | grep -A 20 "goroutine"
Pro Tip: Inspect containerd Daemon Thread Dumps and Tracing
Advertisement
Step 3

Triage Hanging containerd-shim-runc-v2 Processes

A frozen container or broken storage volume can cause a specific `containerd-shim-runc-v2` process to lock up the containerd event monitor. Identify orphaned or deadlocked shims and inspect their wait state in `/proc`.

# Find shims without active runc parent or in 'D' state (uninterruptible sleep)
ps -eo pid,ppid,stat,wchan:20,comm | grep -E 'containerd-shim|D'

# Trace hanging system calls on the containerd daemon
strace -p $(pgrep -x containerd) -f -e trace=network,futex,epoll_wait -s 256
Pro Tip: Triage Hanging containerd-shim-runc-v2 Processes
Step 4

Repair Socket and Safely Recover Without Pod Eviction

If containerd's state database (`/var/lib/containerd/io.containerd.metadata.v1.bolt/meta.db`) or containerd socket is deadlocked, restart containerd. Because containerd uses decoupled shims (`shim-v2`), running containers continue executing uninterrupted during daemon restart.

# Restart containerd daemon (running containers remain alive thanks to shims)
systemctl restart containerd

# Verify CRI endpoint responsiveness
crictl --runtime-endpoint unix:///run/containerd/containerd.sock info | jq .status
Pro Tip: Repair Socket and Safely Recover Without Pod Eviction
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"containerd decouples the runtime daemon from running containers using `containerd-shim-runc-v2`. When the CRI socket hangs, `crictl` hangs because it uses the CRI gRPC plugin, whereas `ctr` can verify if containerd's core runtime is functional. Restarting containerd preserves running workloads."
⚡ 60-Second Elevator Pitch Talking Points
  • W
  • h
  • e
  • n
  • a
  • n
  • o
  • d
  • e
  • e
  • n
  • t
  • e
  • r
  • s
  • N
  • o
  • t
  • R
  • e
  • a
  • d
  • y
  • d
  • u
  • e
  • t
  • o
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • d
  • C
  • R
  • I
  • s
  • o
  • c
  • k
  • e
  • t
  • t
  • i
  • m
  • e
  • o
  • u
  • t
  • s
  • ,
  • w
  • e
  • s
  • e
  • n
  • d
  • `
  • S
  • I
  • G
  • U
  • S
  • R
  • 1
  • `
  • t
  • o
  • c
  • a
  • p
  • t
  • u
  • r
  • e
  • g
  • o
  • r
  • o
  • u
  • t
  • i
  • n
  • e
  • d
  • e
  • a
  • d
  • l
  • o
  • c
  • k
  • s
  • a
  • n
  • d
  • i
  • d
  • e
  • n
  • t
  • i
  • f
  • y
  • b
  • l
  • o
  • c
  • k
  • e
  • d
  • s
  • h
  • i
  • m
  • s
  • i
  • n
  • u
  • n
  • i
  • n
  • t
  • e
  • r
  • r
  • u
  • p
  • t
  • i
  • b
  • l
  • e
  • s
  • l
  • e
  • e
  • p
  • (
  • '
  • D
  • '
  • s
  • t
  • a
  • t
  • e
  • )
  • .
  • B
  • e
  • c
  • a
  • u
  • s
  • e
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • d
  • d
  • e
  • l
  • e
  • g
  • a
  • t
  • e
  • s
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • l
  • i
  • f
  • e
  • c
  • y
  • c
  • l
  • e
  • s
  • t
  • o
  • d
  • e
  • c
  • o
  • u
  • p
  • l
  • e
  • d
  • s
  • h
  • i
  • m
  • p
  • r
  • o
  • c
  • e
  • s
  • s
  • e
  • s
  • ,
  • r
  • e
  • s
  • t
  • a
  • r
  • t
  • i
  • n
  • g
  • `
  • c
  • o
  • n
  • t
  • a
  • i
  • n
  • e
  • r
  • d
  • `
  • r
  • e
  • c
  • o
  • n
  • n
  • e
  • c
  • t
  • s
  • t
  • h
  • e
  • C
  • R
  • I
  • s
  • o
  • c
  • k
  • e
  • t
  • t
  • o
  • e
  • x
  • i
  • s
  • t
  • i
  • n
  • g
  • p
  • o
  • d
  • s
  • w
  • i
  • t
  • h
  • o
  • u
  • t
  • t
  • e
  • r
  • m
  • i
  • n
  • a
  • t
  • i
  • n
  • g
  • c
  • u
  • s
  • t
  • o
  • m
  • e
  • r
  • w
  • o
  • r
  • k
  • l
  • o
  • a
  • d
  • s
  • .
Advertisement
Want more Docker & Containers scenarios?
Explore our complete collection of scenario-based Docker & Containers interview runbooks.
Browse All Docker & Containers Questions →