The Complete Kubernetes Interview Handbook: 100 Questions, Real-World Incidents & Architecture Mastery
DevOps interviews are no longer about reciting textbook definitions or memorizing CLI flags. Top engineering teams evaluate your incident triage instincts: when a pod is in CrashLoopBackOff, the Jenkins pipeline is green but the service is returning 503s, or node CPU looks normal while users experience latency spikes.
Module 1: Basics & Architecture (Q1–Q10)
In Senior and Staff-level interviews, questions about Kubernetes architecture do not ask you to list components; they test what happens when components fail.
What happens when the Control Plane components fail?
- kube-apiserver down: Existing running pods continue uninterrupted. However, no new deployments, scaling operations, or secret changes can occur.
kubectlwill return connection refused. - etcd quorum loss: When a 3-node etcd cluster loses 2 nodes, write operations halt immediately (read-only mode). The API server cannot persist new object states.
- kube-scheduler down: Newly created pods remain in
Pendingstate indefinitely. Existing pods continue running on assigned worker nodes. - kube-controller-manager down: Replicas will not self-heal if a node dies, and rolling updates freeze midway.
Critical etcd Health & Backup Verification Commands
# Check etcd member health & quorum status
ETCDCTL_API=3 etcdctl --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key --endpoints=https://127.0.0.1:2379 endpoint health -w table
# Create consistent snapshot backup
ETCDCTL_API=3 etcdctl --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/server.crt --key=/etc/kubernetes/pki/etcd/server.key --endpoints=https://127.0.0.1:2379 snapshot save /backup/etcd-snapshot-$(date +%Y%m%d).db
# Verify snapshot integrity before restoring
ETCDCTL_API=3 etcdctl snapshot status /backup/etcd-snapshot-$(date +%Y%m%d).db -w table
Module 2: Pods & Workload Lifecycle (Q11–Q22)
Modern Kubernetes architectures heavily utilize multi-container patterns: Init Containers (schema migrations, pre-flight dependency checks), Native Sidecar Containers (Kubernetes 1.28+ restartPolicy: Always in initContainers), and Ephemeral Debug Containers for distroless, minimal images.
Debugging Distroless Pods with Ephemeral Containers
When your production container has no shell, curl, or netstat (e.g., Google Distroless or Chainguard images), use kubectl debug to attach a live diagnostic container sharing the pod's network and process namespaces:
# Attach netshoot ephemeral container to live failing pod
kubectl debug -it pod/payment-service-784d858cf8-xkp82 --image=nicolaka/netshoot --target=payment-app -- /bin/bash
# Inside netshoot: test TCP socket connectivity and DNS latency
curl -v -m 2 http://order-db.internal:5432/
dig +trace +stats order-db.internal
Module 3: Services & Cloud Networking (Q23–Q34)
Kubernetes service discovery relies on virtual IP routing (iptables, IPVS, or eBPF with Cilium) and CoreDNS.
The Infamous 5-Second DNS Latency Bug (ndots:5 & UDP Conntrack Race)
By default, Kubernetes injects ndots:5 into /etc/resolv.conf. When querying external domains like api.stripe.com (which has fewer than 5 dots), the resolver appends search paths sequentially (api.stripe.com.default.svc.cluster.local, etc.), firing 5 unnecessary DNS queries per connection. Additionally, glibc concurrently sends A and AAAA records over UDP from the same socket, causing Linux conntrack race conditions that result in a 5000ms timeout!
dnsConfig.options with ndots:2 and single-request-reopen.
Module 4: Configuration & Persistent Storage (Q35–Q44)
Storage in cloud-native environments requires decoupling storage implementation from application logic using StorageClasses, Dynamic Volume Provisioning, and CSI plugins.
- VolumeBindingMode: WaitForFirstConsumer: Crucial in multi-AZ clusters (AWS EBS, GCP Persistent Disk) to ensure the PV is created in the exact Availability Zone where the pod is scheduled.
- ConfigMap Hot Reloads: ConfigMaps mounted via volumes are automatically updated via symlink rotation; ConfigMaps injected via environment variables (
envFrom) require a pod restart. Use Reloader controller for automated rolling updates. - External Secrets Operator (ESO): Syncs production credentials directly from AWS Secrets Manager or HashiCorp Vault into native Kubernetes Secrets without storing plaintext in Git.
Module 5: Scheduling & Resource Management (Q45–Q54)
The Kubernetes scheduler operates in two main phases: Filtering (Predicates) and Scoring (Priorities). Understanding CPU CFS quota throttling and Memory OOM behavior is essential.
CPU Throttling vs Memory OOMKill (Exit Code 137)
CPU is a compressible resource: when a container hits its CPU limit, the Linux kernel throttles CFS periods, causing slow response times without terminating the process. Memory is an incompressible resource: when a container exceeds its memory limit, the Linux kernel OOM Killer immediately kills the highest-consuming process with Exit Code 137 (128 + SIGKILL 9).
Module 6: Scaling, Availability & Disruption (Q55–Q62)
High availability requires combining Horizontal Pod Autoscaler (HPA), PodDisruptionBudgets (PDB), and graceful pod termination.
# High Availability PodDisruptionBudget YAML
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payment-service-pdb
namespace: production
spec:
minAvailable: 75%
selector:
matchLabels:
app: payment-service
Always configure preStop hooks with a small sleep (e.g. sleep 15) so that Ingress and Service endpoints remove the pod from routing tables before the container receives SIGTERM, eliminating in-flight 502/503 errors during rolling updates.
Module 7: Security, RBAC & Pod Standards (Q63–Q70)
Zero-Trust cluster architecture requires combining Role-Based Access Control (RBAC), TokenRequest API, default-deny NetworkPolicies, and Pod Security Admission (Restricted profile).
# Default Deny All Ingress & Egress NetworkPolicy
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
Module 8: Observability, Metrics & Telemetry (Q71–Q75)
Telemetry in modern clusters spans three distinct layers:
- Cluster State:
kube-state-metricsexports object health (deployment replicas desired vs available, pending pods). - Container Resources:
cAdvisorembedded in kubelet exports CPU, memory, filesystem, and network socket consumption. - Application Metrics: Prometheus
ServiceMonitorandPodMonitorCRDs scraping/metricsendpoints with OpenTelemetry distributed traces.
Module 9: CI/CD Workflows & Helm Delivery (Q76–Q80)
Helm manages application releases with templating, hooks, and release rollback mechanisms. When paired with GitOps engines like ArgoCD or FluxCD, the cluster state is continuously reconciled against a Git single source of truth.
# Check Helm release status and rollback safely
helm history checkout-service -n production
helm rollback checkout-service 4 -n production
Pod is Running. Node is Healthy. CPU & Memory are Normal. But Users Get 503s!
Every DevOps/SRE engineer has faced this in production. And it is one of the most frequently asked interview questions: The truth: Pod Running ≠ Application Healthy.
The Step-by-Step 9-Stage Troubleshooting Sequence:
- 1️⃣ Check Pod Status: Run
kubectl get pods -o wideandkubectl describe pod <pod-name>to verify if restarts occurred. - 2️⃣ Inspect Previous Crash Logs: Use
kubectl logs <pod> --previousto see if the process exited and restarted. - 3️⃣ Review Real-Time Events: Check for probe failures, scheduling errors, or container restarts using
kubectl get events --sort-by=.metadata.creationTimestamp. - 4️⃣ Verify Service Endpoints: Check if pods are attached to endpoints:
kubectl get endpoints <svc-name>. If endpoints list is empty, pod labels do not match the service selector! - 5️⃣ Check Ingress / Load Balancer Rules: Ensure target port matches container port and health check paths align with ingress annotations.
- 6️⃣ Verify Application Readiness Probe: If
readinessProbeis failing, the pod remains inRunningstate but is immediately removed from Endpoint routing, resulting in 503 Service Unavailable! - 7️⃣ Validate ConfigMaps & Secrets: Verify injected environment variables and secret credentials with
kubectl exec -it <pod> -- printenv. - 8️⃣ Test Upstream Dependencies: Exec into pod to check connection to PostgreSQL, Redis, Kafka, or external payment gateways:
nc -zv db.internal 5432. - 9️⃣ Inspect Recent Deployment History: Check rollout history with
kubectl rollout history deployment/<name>to spot recent configuration commits.
🔥 Most Common Root Causes for 503 Errors:
- Readiness probe failing, preventing the pod from receiving traffic.
- Selector mismatch between Service and Deployment labels.
- Downstream database or Redis connection pool exhaustion.
- Ingress controller upstream keep-alive timeout shorter than backend application timeout.
DevOps Interviews Are No Longer About Commands — Master Real Scenarios
Modern interviews test your operational maturity under real production stress. Here are the core question patterns and how to formulate winning STAR responses:
1. "Jenkins Pipeline is GREEN, but App is DOWN"
Investigate if the CI pipeline pushed the image and updated the manifest, but the CD agent encountered an ImagePullBackOff (bad ECR credentials or typo in SHA tag), or the app booted and failed health checks immediately after the deployment stage completed.
2. "Users Report Slowness, but CPU/RAM Normal"
Check thread locks, database connection pool exhaustion (hikari/pgBouncer waiting queues), external 3rd-party API response timeouts, or TCP socket exhaustion (TIME_WAIT states on the host network).
3. "A Node Becomes NotReady in Production"
Inspect kubelet logs on the node (journalctl -u kubelet), check DiskPressure/MemoryPressure conditions, verify container runtime health (containerd socket), and monitor pod eviction grace periods.
4. "Deployment Succeeded, but Traffic Dropping"
Check if an environment variable changed between staging and production (e.g. CORS origin, DB host), verify SSL certificates, and check ALB 504 gateway timeout alerts.
Get the Full 101-Page Kubernetes Interview Handbook PDF
Take all 100 questions, diagnostic command trees, incident resolution checklists, and architecture diagrams with you offline. Free instant download, no email registration required.