⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Docker Must enable BuildKit Staff SRE Scenario [L3]

Q: Your team wants to implement live migration of a running Docker container from one host to another without stopping it, similar to VM live migration. Is this possible with Docker? What technology enables it?

Docker has experimental support for checkpoint and restore using CRIU (Checkpoint/Restore In Userspace). CRIU freezes a running process, ...

#Docker #Must enable BuildKit #L3 #Containers #Linux #Terraform State
🎙️ Candidate Opening & Architectural Context
""Container stability relies on clean signal handling (SIGTERM vs SIGKILL) and immutable image tagging. The interviewer is testing: CRIU (Checkpoint/Restore in Userspace), container migration limitations.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Docker has experimental support for checkpoint and restore using CRIU (Checkpoint/Restore In Userspace). CRIU freezes a running process, serializes its entire state (memory, registers, open files, sockets, timers) to disk, and can restore it later — even on a different host.

  • Checkpoint: docker checkpoint create checkpoint1
  • Transfer the checkpoint data and container filesystem to the target host.
  • Restore: docker start --checkpoint checkpoint1
2️⃣

Remediation & Permanent Safeguards

Workflow: *Limitations:* This feature is experimental and not production-ready. Open network connections break (TCP state doesn't survive cross-host migration). External storage must be shared (e.g., NFS). GPU state, complex IPC, and certain kernel features aren't fully supported. For production workloads, Kubernetes pod rescheduling with graceful shutdown/startup is the practical alternative.

💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Checkpoint: docker checkpoint create checkpoint1."
⚡ 60-Second Elevator Pitch Talking Points
  • Checkpoint: docker checkpoint create checkpoint1
  • Transfer the checkpoint data and container filesystem to the target host.
  • Restore: docker start --checkpoint checkpoint1
Advertisement
Want more Docker scenarios?
Explore our complete collection of scenario-based Docker interview runbooks.
Browse All Docker Questions →

📚 Related Production Scenarios in Docker