Q: Your SRE team must apply an urgent security patch to `dockerd` across a fleet of 50 standalone production hosts running stateful databases and long-running analytics jobs. Historically, running `systemctl restart docker` instantly killed all running containers on the host, forcing scheduled weekend maintenance windows. You must configure and validate Docker's `live-restore` feature, ensuring that containers remain running, network connectivity stays open, and the daemon reconnects seamlessly upon restart.
Configure Docker's `live-restore` capability to perform engine upgrades, daemon patching, and configuration reloads without terminating running container workloads.
Want to master this scenario in a live sandbox? KodeKloud's Docker Certified Associate (DCA) Hands-On Lab Course covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Understand containerd-shim Decoupling and Live-Restore Mechanics
By default, when `dockerd` shuts down, it terminates all child processes. Enabling `live-restore` instructs `dockerd` to leave the `containerd-shim` processes running. The shims maintain stdout/stderr FIFOs, cgroups, and network namespaces for active containers while `dockerd` restarts and reconnects to the shim sockets.
<!-- Live-Restore Decoupled Architecture -->
During dockerd Restart:
dockerd (Stopping / Upgrading / Starting)
✗ (Disconnected temporarily)
containerd-shim-v2 ───> Container App (RUNNING UNINTERRUPTED!)
containerd-shim-v2 ───> Container App (RUNNING UNINTERRUPTED!)
Enable live-restore in /etc/docker/daemon.json
Add `"live-restore": true` to the Docker daemon configuration and apply via `SIGHUP` reload.
cat <<EOF | sudo tee /etc/docker/daemon.json
{
"live-restore": true
}
EOF
# Reload daemon configuration without restarting
sudo systemctl reload docker
# Verify live-restore is enabled
docker info --format '{{.LiveRestoreEnabled}}'
# Output: true
Execute and Verify Daemon Restart During Active Traffic
Spawn a long-running container streaming telemetry or network packets. Restart the Docker daemon service and observe that the container process maintains 100% uptime.
# Run container logging timestamps
docker run -d --name uptime-test alpine sh -c 'while true; do date; sleep 1; done'
# Restart docker service
sudo systemctl restart docker
# Verify container PID was never interrupted
docker inspect uptime-test --format '{{.State.Status}} (PID: {{.State.Pid}})'
# Output confirms status is 'running' with identical host PID!
Review live-restore Constraints and Limitations
Be aware of operational boundaries: `live-restore` is not supported in Docker Swarm mode (Swarm manager assumes node failure if daemon disconnects), and daemon major version upgrades that modify the shim IPC protocol require container recreation.
- W
- e
- e
- l
- i
- m
- i
- n
- a
- t
- e
- d
- s
- c
- h
- e
- d
- u
- l
- e
- d
- m
- a
- i
- n
- t
- e
- n
- a
- n
- c
- e
- d
- o
- w
- n
- t
- i
- m
- e
- f
- o
- r
- h
- o
- s
- t
- e
- n
- g
- i
- n
- e
- p
- a
- t
- c
- h
- i
- n
- g
- b
- y
- e
- n
- a
- b
- l
- i
- n
- g
- D
- o
- c
- k
- e
- r
- `
- l
- i
- v
- e
- -
- r
- e
- s
- t
- o
- r
- e
- `
- .
- B
- e
- c
- a
- u
- s
- e
- c
- o
- n
- t
- a
- i
- n
- e
- r
- d
- s
- h
- i
- m
- s
- m
- a
- i
- n
- t
- a
- i
- n
- c
- o
- n
- t
- a
- i
- n
- e
- r
- l
- i
- f
- e
- c
- y
- c
- l
- e
- s
- i
- n
- d
- e
- p
- e
- n
- d
- e
- n
- t
- l
- y
- o
- f
- t
- h
- e
- m
- a
- i
- n
- d
- a
- e
- m
- o
- n
- ,
- w
- e
- c
- a
- n
- r
- e
- s
- t
- a
- r
- t
- a
- n
- d
- u
- p
- g
- r
- a
- d
- e
- `
- d
- o
- c
- k
- e
- r
- d
- `
- d
- u
- r
- i
- n
- g
- b
- u
- s
- i
- n
- e
- s
- s
- h
- o
- u
- r
- s
- w
- i
- t
- h
- o
- u
- t
- d
- r
- o
- p
- p
- i
- n
- g
- a
- s
- i
- n
- g
- l
- e
- a
- c
- t
- i
- v
- e
- c
- u
- s
- t
- o
- m
- e
- r
- s
- o
- c
- k
- e
- t
- o
- r
- r
- e
- s
- t
- a
- r
- t
- i
- n
- g
- s
- t
- a
- t
- e
- f
- u
- l
- p
- r
- o
- c
- e
- s
- s
- e
- s
- .