Q: How would you ensure a critical pod always runs on the same node?
Two approaches:
#Kubernetes #Deployments & Workloads #L2 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""When troubleshooting Kubernetes, I always follow a structured layered model: Pod status -> Events -> Logs -> Network. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
Two approaches:
- NodeSelector — add a label to the node (
kubectl label node) and addtype=critical nodeSelector: {type: critical}to the pod spec. Simple but inflexible. - Node Affinity — more expressive, supports
requiredDuringSchedulingIgnoredDuringExecution(hard rule) orpreferredDuringScheduling...(soft preference).
2️⃣
Remediation & Permanent Safeguards
For "always the same node" — use requiredDuringSchedulingIgnoredDuringExecution with nodeAffinity. But be careful: if that node goes down, the pod won't reschedule elsewhere.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: NodeSelector — add a label to the node (kubectl label node type=critical) and add nodeSelector: {type: critical} to the pod spec.."
⚡ 60-Second Elevator Pitch Talking Points
- NodeSelector — add a label to the node (kubectl label node type=critical) and add nodeSelector: ...
- Node Affinity — more expressive, supports requiredDuringSchedulingIgnoredDuringExecution (hard ru...
Advertisement