⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
☸️ Flagship SRE Architecture Blueprint · 101-Page Companion

The Complete Kubernetes Interview Handbook: 100 Questions, Real-World Incidents & Architecture Mastery

DevOps interviews are no longer about reciting textbook definitions or memorizing CLI flags. Top engineering teams evaluate your incident triage instincts: when a pod is in CrashLoopBackOff, the Jenkins pipeline is green but the service is returning 503s, or node CPU looks normal while users experience latency spikes.

📥 Free 101-Page Companion PDF Handbook
Download the Complete Offline Handbook
Includes all 100 questions, printable diagnostic trees, kubectl cheat sheets, and production YAML templates (209 KB).
📥 Download PDF (Free) 📚 All PDF Guides

Module 1: Basics & Architecture (Q1–Q10)

In Senior and Staff-level interviews, questions about Kubernetes architecture do not ask you to list components; they test what happens when components fail.

What happens when the Control Plane components fail?

  • kube-apiserver down: Existing running pods continue uninterrupted. However, no new deployments, scaling operations, or secret changes can occur. kubectl will return connection refused.
  • etcd quorum loss: When a 3-node etcd cluster loses 2 nodes, write operations halt immediately (read-only mode). The API server cannot persist new object states.
  • kube-scheduler down: Newly created pods remain in Pending state indefinitely. Existing pods continue running on assigned worker nodes.
  • kube-controller-manager down: Replicas will not self-heal if a node dies, and rolling updates freeze midway.

Critical etcd Health & Backup Verification Commands

# Check etcd member health & quorum status
ETCDCTL_API=3 etcdctl --cacert=/etc/kubernetes/pki/etcd/ca.crt   --cert=/etc/kubernetes/pki/etcd/server.crt   --key=/etc/kubernetes/pki/etcd/server.key   --endpoints=https://127.0.0.1:2379 endpoint health -w table

# Create consistent snapshot backup
ETCDCTL_API=3 etcdctl --cacert=/etc/kubernetes/pki/etcd/ca.crt   --cert=/etc/kubernetes/pki/etcd/server.crt   --key=/etc/kubernetes/pki/etcd/server.key   --endpoints=https://127.0.0.1:2379 snapshot save /backup/etcd-snapshot-$(date +%Y%m%d).db

# Verify snapshot integrity before restoring
ETCDCTL_API=3 etcdctl snapshot status /backup/etcd-snapshot-$(date +%Y%m%d).db -w table
        

Module 2: Pods & Workload Lifecycle (Q11–Q22)

Modern Kubernetes architectures heavily utilize multi-container patterns: Init Containers (schema migrations, pre-flight dependency checks), Native Sidecar Containers (Kubernetes 1.28+ restartPolicy: Always in initContainers), and Ephemeral Debug Containers for distroless, minimal images.

Debugging Distroless Pods with Ephemeral Containers

When your production container has no shell, curl, or netstat (e.g., Google Distroless or Chainguard images), use kubectl debug to attach a live diagnostic container sharing the pod's network and process namespaces:

# Attach netshoot ephemeral container to live failing pod
kubectl debug -it pod/payment-service-784d858cf8-xkp82   --image=nicolaka/netshoot   --target=payment-app   -- /bin/bash

# Inside netshoot: test TCP socket connectivity and DNS latency
curl -v -m 2 http://order-db.internal:5432/
dig +trace +stats order-db.internal
          

Module 3: Services & Cloud Networking (Q23–Q34)

Kubernetes service discovery relies on virtual IP routing (iptables, IPVS, or eBPF with Cilium) and CoreDNS.

The Infamous 5-Second DNS Latency Bug (ndots:5 & UDP Conntrack Race)

By default, Kubernetes injects ndots:5 into /etc/resolv.conf. When querying external domains like api.stripe.com (which has fewer than 5 dots), the resolver appends search paths sequentially (api.stripe.com.default.svc.cluster.local, etc.), firing 5 unnecessary DNS queries per connection. Additionally, glibc concurrently sends A and AAAA records over UDP from the same socket, causing Linux conntrack race conditions that result in a 5000ms timeout!

Senior SRE Solution: Deploy NodeLocal DNSCache or tune dnsConfig.options with ndots:2 and single-request-reopen.

Module 4: Configuration & Persistent Storage (Q35–Q44)

Storage in cloud-native environments requires decoupling storage implementation from application logic using StorageClasses, Dynamic Volume Provisioning, and CSI plugins.

  • VolumeBindingMode: WaitForFirstConsumer: Crucial in multi-AZ clusters (AWS EBS, GCP Persistent Disk) to ensure the PV is created in the exact Availability Zone where the pod is scheduled.
  • ConfigMap Hot Reloads: ConfigMaps mounted via volumes are automatically updated via symlink rotation; ConfigMaps injected via environment variables (envFrom) require a pod restart. Use Reloader controller for automated rolling updates.
  • External Secrets Operator (ESO): Syncs production credentials directly from AWS Secrets Manager or HashiCorp Vault into native Kubernetes Secrets without storing plaintext in Git.

Module 5: Scheduling & Resource Management (Q45–Q54)

The Kubernetes scheduler operates in two main phases: Filtering (Predicates) and Scoring (Priorities). Understanding CPU CFS quota throttling and Memory OOM behavior is essential.

CPU Throttling vs Memory OOMKill (Exit Code 137)

CPU is a compressible resource: when a container hits its CPU limit, the Linux kernel throttles CFS periods, causing slow response times without terminating the process. Memory is an incompressible resource: when a container exceeds its memory limit, the Linux kernel OOM Killer immediately kills the highest-consuming process with Exit Code 137 (128 + SIGKILL 9).

Module 6: Scaling, Availability & Disruption (Q55–Q62)

High availability requires combining Horizontal Pod Autoscaler (HPA), PodDisruptionBudgets (PDB), and graceful pod termination.

# High Availability PodDisruptionBudget YAML
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: payment-service-pdb
  namespace: production
spec:
  minAvailable: 75%
  selector:
    matchLabels:
      app: payment-service
        

Always configure preStop hooks with a small sleep (e.g. sleep 15) so that Ingress and Service endpoints remove the pod from routing tables before the container receives SIGTERM, eliminating in-flight 502/503 errors during rolling updates.

Module 7: Security, RBAC & Pod Standards (Q63–Q70)

Zero-Trust cluster architecture requires combining Role-Based Access Control (RBAC), TokenRequest API, default-deny NetworkPolicies, and Pod Security Admission (Restricted profile).

# Default Deny All Ingress & Egress NetworkPolicy
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-all
  namespace: production
spec:
  podSelector: {}
  policyTypes:
  - Ingress
  - Egress
        

Module 8: Observability, Metrics & Telemetry (Q71–Q75)

Telemetry in modern clusters spans three distinct layers:

  1. Cluster State: kube-state-metrics exports object health (deployment replicas desired vs available, pending pods).
  2. Container Resources: cAdvisor embedded in kubelet exports CPU, memory, filesystem, and network socket consumption.
  3. Application Metrics: Prometheus ServiceMonitor and PodMonitor CRDs scraping /metrics endpoints with OpenTelemetry distributed traces.

Module 9: CI/CD Workflows & Helm Delivery (Q76–Q80)

Helm manages application releases with templating, hooks, and release rollback mechanisms. When paired with GitOps engines like ArgoCD or FluxCD, the cluster state is continuously reconciled against a Git single source of truth.

# Check Helm release status and rollback safely
helm history checkout-service -n production
helm rollback checkout-service 4 -n production
        
🚨 SRE Incident Triage Runbook (Q81–Q90)

Pod is Running. Node is Healthy. CPU & Memory are Normal. But Users Get 503s!

Every DevOps/SRE engineer has faced this in production. And it is one of the most frequently asked interview questions: The truth: Pod Running ≠ Application Healthy.

The Step-by-Step 9-Stage Troubleshooting Sequence:

  1. 1️⃣ Check Pod Status: Run kubectl get pods -o wide and kubectl describe pod <pod-name> to verify if restarts occurred.
  2. 2️⃣ Inspect Previous Crash Logs: Use kubectl logs <pod> --previous to see if the process exited and restarted.
  3. 3️⃣ Review Real-Time Events: Check for probe failures, scheduling errors, or container restarts using kubectl get events --sort-by=.metadata.creationTimestamp.
  4. 4️⃣ Verify Service Endpoints: Check if pods are attached to endpoints: kubectl get endpoints <svc-name>. If endpoints list is empty, pod labels do not match the service selector!
  5. 5️⃣ Check Ingress / Load Balancer Rules: Ensure target port matches container port and health check paths align with ingress annotations.
  6. 6️⃣ Verify Application Readiness Probe: If readinessProbe is failing, the pod remains in Running state but is immediately removed from Endpoint routing, resulting in 503 Service Unavailable!
  7. 7️⃣ Validate ConfigMaps & Secrets: Verify injected environment variables and secret credentials with kubectl exec -it <pod> -- printenv.
  8. 8️⃣ Test Upstream Dependencies: Exec into pod to check connection to PostgreSQL, Redis, Kafka, or external payment gateways: nc -zv db.internal 5432.
  9. 9️⃣ Inspect Recent Deployment History: Check rollout history with kubectl rollout history deployment/<name> to spot recent configuration commits.

🔥 Most Common Root Causes for 503 Errors:

  • Readiness probe failing, preventing the pod from receiving traffic.
  • Selector mismatch between Service and Deployment labels.
  • Downstream database or Redis connection pool exhaustion.
  • Ingress controller upstream keep-alive timeout shorter than backend application timeout.
🎯 Interview Pro-Tip: Never just say "the pod is running". Walk through the full request flow from start to finish: User → DNS Resolution → Cloud LB → Ingress Controller → Service Routing → Endpoints → Pod → Backend Dependencies
🎯 Real Production Scenarios (Q91–Q100)

DevOps Interviews Are No Longer About Commands — Master Real Scenarios

Modern interviews test your operational maturity under real production stress. Here are the core question patterns and how to formulate winning STAR responses:

1. "Jenkins Pipeline is GREEN, but App is DOWN"

Investigate if the CI pipeline pushed the image and updated the manifest, but the CD agent encountered an ImagePullBackOff (bad ECR credentials or typo in SHA tag), or the app booted and failed health checks immediately after the deployment stage completed.

2. "Users Report Slowness, but CPU/RAM Normal"

Check thread locks, database connection pool exhaustion (hikari/pgBouncer waiting queues), external 3rd-party API response timeouts, or TCP socket exhaustion (TIME_WAIT states on the host network).

3. "A Node Becomes NotReady in Production"

Inspect kubelet logs on the node (journalctl -u kubelet), check DiskPressure/MemoryPressure conditions, verify container runtime health (containerd socket), and monitor pod eviction grace periods.

4. "Deployment Succeeded, but Traffic Dropping"

Check if an environment variable changed between staging and production (e.g. CORS origin, DB host), verify SSL certificates, and check ALB 504 gateway timeout alerts.

☸️ Practice 175+ Live Kubernetes Scenarios
📥

Get the Full 101-Page Kubernetes Interview Handbook PDF

Take all 100 questions, diagnostic command trees, incident resolution checklists, and architecture diagrams with you offline. Free instant download, no email registration required.

📥 Download PDF (209 KB · 101 Pages) 📚 View All Free PDF Resources
Advertisement