Q: What is the role of etcd in Kubernetes and what happens if etcd goes down?
etcd is the key-value store that is Kubernetes' "brain." Every cluster state (pod specs, node info, secrets, configmaps, events) is store...
#Kubernetes #Advanced Scenarios #L3 #Container Orchestration #K8s #etcd
🎙️ Candidate Opening & Architectural Context
""In one of our high-traffic production clusters, our SRE team handled this incident using a standardized runbook. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
etcd is the key-value store that is Kubernetes' "brain." Every cluster state (pod specs, node info, secrets, configmaps, events) is stored in etcd.
- Existing pods keep running — kubelet runs pods independently of the API server.
- No new pods can be created — API server can't write new state.
- No changes work — no scaling, no new deployments, no config changes.
2️⃣
Remediation & Permanent Safeguards
If etcd goes down: Recovery: restore etcd from a snapshot backup. This is why etcd backups are critical (using etcdctl snapshot save). Production setup: etcd should have an odd number of nodes (3, 5) for quorum. With 3 nodes, cluster can tolerate 1 failure. With 5 nodes, 2 failures.
- The cluster is effectively read-only and frozen.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Existing pods keep running — kubelet runs pods independently of the API server.."
⚡ 60-Second Elevator Pitch Talking Points
- Existing pods keep running — kubelet runs pods independently of the API server.
- No new pods can be created — API server can't write new state.
- No changes work — no scaling, no new deployments, no config changes.
Advertisement