Q: What are the most critical Kubernetes interview questions asked in senior DevOps and SRE interview loops? Walk me through the core architecture, pod lifecycle, service networking, and troubleshooting methodologies.
The definitive master preparation guide for Kubernetes interview questions and answers: control plane internals (etcd, API server, scheduler), pod lifecycle states, CNI networking, ingress routing, and production incident triage.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Pillar 1: Control Plane Mechanics & etcd Quorum
How the Kubernetes brain coordinates cluster state:
- kube-apiserver: The sole component that communicates with etcd. Validates, authenticates, mutates, and admits requests via admission webhooks.
- etcd Quorum & Raft: Distributed key-value store requiring an odd number of members (2N + 1). A 3-node cluster tolerates 1 node failure; a 5-node cluster tolerates 2 node failures.
- kube-scheduler: Two-phase scheduling algorithm: Filtering (predicates - finding nodes with resources/tolerations) and Scoring (priorities - ranking best nodes based on image locality and spread constraints).
- kube-controller-manager: Runs core reconciliation loops (DeploymentController, NodeController, EndpointsController) continuously driving actual state to desired state.
Pillar 2: Pod Lifecycle & Health Probes
Understanding container initiation, health checks, and graceful shutdown:
- Startup Probe: Disables liveness and readiness checks during slow application boots (e.g. JVM warmup). If it fails after failureThreshold, container restarts.
- Readiness Probe: Determines if pod should receive traffic via Service endpoints. Failure removes pod from endpoint slices without restarting container.
- Liveness Probe: Detects deadlocks. Failure triggers Kubelet container kill and restart according to restartPolicy.
- Graceful Termination: Pod marked Terminating → removed from endpoints → preStop hook executes → SIGTERM sent → terminationGracePeriodSeconds (default 30s) → SIGKILL.
Pillar 3: Kubernetes Networking & Service Discovery
How packets flow across pods, services, and external ingress:
- Fundamental Network Model: Every pod receives a unique IP; all pods communicate without NAT; node-to-pod communication is flat.
- kube-proxy vs Cilium eBPF: Legacy kube-proxy translates Service ClusterIPs via iptables or IPVS rules (O(N) lookup degradation at high service counts). Cilium replaces iptables with eBPF maps for O(1) packet translation.
- CoreDNS Resolution: Queries follow the pattern
. Beware the ndots:5 latency bug causing multiple trailing domain lookups.. .svc.cluster.local
Pillar 4: Production Incident Triage Runbook
The four classic Kubernetes production failures and their systematic diagnosis:
- CrashLoopBackOff: Check
kubectl logsto inspect why previous container crashed, followed by--previous kubectl describe pod. - ImagePullBackOff: Inspect pod Events for registry authentication errors, tag typos, or Docker Hub 429 rate limit errors.
- OOMKilled (Exit Code 137): Container exceeded cgroup memory limit; inspect
kubectl describe podLast State: Terminated Reason: OOMKilled and increase memory limits. - Pending Pods: Scheduler cannot find a suitable node; check Events for Insufficient cpu/memory, untolerated node taints, or volume affinity conflicts.
- Articulated Kubernetes control plane internals, etcd Raft quorum requirements, and scheduler filter/score phases.
- Architected graceful pod lifecycles combining startup/readiness/liveness probes with preStop hooks.
- Compared kube-proxy iptables with modern Cilium eBPF datapaths for high-scale microservice networking.
- Demonstrated systematic terminal debugging for CrashLoopBackOff, ImagePullBackOff, OOMKilled, and Pending pods.