⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Kubernetes Advanced Scenarios Staff SRE Scenario [L3]

Q: What is the role of etcd in Kubernetes and what happens if etcd goes down?

etcd is the key-value store that is Kubernetes' "brain." Every cluster state (pod specs, node info, secrets, configmaps, events) is store...

#Kubernetes #Advanced Scenarios #L3 #Container Orchestration #K8s #etcd
🎙️ Candidate Opening & Architectural Context
""In one of our high-traffic production clusters, our SRE team handled this incident using a standardized runbook. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

etcd is the key-value store that is Kubernetes' "brain." Every cluster state (pod specs, node info, secrets, configmaps, events) is stored in etcd.

  • Existing pods keep running — kubelet runs pods independently of the API server.
  • No new pods can be created — API server can't write new state.
  • No changes work — no scaling, no new deployments, no config changes.
2️⃣

Remediation & Permanent Safeguards

If etcd goes down: Recovery: restore etcd from a snapshot backup. This is why etcd backups are critical (using etcdctl snapshot save). Production setup: etcd should have an odd number of nodes (3, 5) for quorum. With 3 nodes, cluster can tolerate 1 failure. With 5 nodes, 2 failures.

  • The cluster is effectively read-only and frozen.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Existing pods keep running — kubelet runs pods independently of the API server.."
⚡ 60-Second Elevator Pitch Talking Points
  • Existing pods keep running — kubelet runs pods independently of the API server.
  • No new pods can be created — API server can't write new state.
  • No changes work — no scaling, no new deployments, no config changes.
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes