⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff / Principal SRE Kubernetes Cluster Reliability & Diagnostics Core Diagnostics

Q: Your Kubernetes cluster is experiencing intermittent pod failures and high latency. How would you troubleshoot it systematically? Explain how you would investigate pods, nodes, networking, resource limits, probes, scheduling, DNS, and application metrics.

Exhaustive diagnostic methodology for intermittent Kubernetes failures: node kernel pressure, CPU throttling, CoreDNS latency (ndots:5), CNI IP exhaustion, and conntrack table drops.

#Kubernetes #CoreDNS #Latency #Troubleshooting #Conntrack #Throttling #CNI
🎙️ Candidate Opening & Architectural Context
"Intermittent failures and latency spikes in Kubernetes are notoriously elusive because they rarely show up as hard crashes. I isolate them using a layered, full-stack diagnostic model: Nodes -> Pods & Limits -> CoreDNS -> Networking/CNI -> Application Probes."
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Layer 1: Node Health, Kernel & CPU Throttling

Inspect underlying host instances and container runtime:

  • Run kubectl get nodes: Check for node conditions like MemoryPressure, DiskPressure, or PIDPressure.
  • SSH into suspected nodes: check dmesg -T for kernel OOM-killer invocations, hardware errors, or TCP drops.
  • CPU Throttling Trap: Check Prometheus metric container_cpu_cfs_throttled_seconds_total. Even when average node CPU is only 40%, strict pod resources.limits.cpu cause Linux CFS throttling, introducing random 200–500ms latency spikes!
2️⃣

Layer 2: Pod Restarts, OOMKills & Exit Codes

Inspect container states across namespaces:

  • kubectl get pods -A --sort-by='.status.containerStatuses[0].restartCount': Identify flapping pods.
  • kubectl describe pod <pod>: Check for OOMKilled (Exit Code 137).
  • Retrieve stack trace from crashed container: kubectl logs <pod> -c <container> --previous.
  • Check liveness probe thresholds: Are slow database calls causing the liveness probe to timeout and restart otherwise healthy pods?
3️⃣

Layer 3: CoreDNS & The 'ndots:5' DNS Latency Trap

Intermittent 1-second latency spikes are almost always DNS issues:

  • Check CoreDNS latency: coredns_dns_request_duration_seconds in Prometheus.
  • The ndots:5 Amplification Trap: Default Kubernetes /etc/resolv.conf has ndots:5. For external domains (e.g. api.stripe.com), the pod queries 4 internal search domains (.default.svc...) before querying the public domain, multiplying DNS queries by 5x and overloading CoreDNS!
  • Fix: Deploy NodeLocal DNSCache daemonset on all nodes, or append a trailing dot (api.stripe.com.) in application configs.
4️⃣

Layer 4: CNI, IP Exhaustion & Conntrack Table

Subtle networking drops at the host and CNI level:

  • VPC CNI IP Exhaustion: In AWS, check if worker node subnets ran out of free private IP addresses, preventing newly scheduled pods from obtaining an ENI secondary IP.
  • Linux Conntrack Saturation: High-traffic microservices exhaust the Linux connection tracking table (nf_conntrack_max). Once full, the kernel silently drops new TCP SYN packets! Check with dmesg -T | grep 'table full, dropping packet'.
  • kube-proxy Sync Latency: Check if iptables rule processing is stalling packet forwarding.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Intermittent K8s latency is usually not pod crashes: it's Linux CFS CPU throttling, CoreDNS ndots:5 search domain multiplication, or Linux nf_conntrack table exhaustion. Deploy NodeLocal DNSCache and tune CPU limits."
⚡ 60-Second Elevator Pitch Talking Points
  • Layer 1 (Node): Check node pressure flags (Memory/DiskPressure) and CFS CPU throttling (container_cpu_cfs_throttled_seconds_total).
  • Layer 2 (Pods): Sort pods by restart count; inspect 'kubectl logs --previous' for OOMKilled (Exit Code 137).
  • Layer 3 (DNS): Inspect CoreDNS metrics; mitigate ndots:5 search domain amplification using NodeLocal DNSCache.
  • Layer 4 (Network): Verify VPC CNI subnet IP availability; check 'dmesg' for nf_conntrack table exhaustion drops.
  • Layer 5 (Probes): Ensure liveness probes have sufficient timeout/initialDelay to avoid killing slow-starting pods.
Advertisement
Want more Kubernetes scenarios?
Explore our complete collection of scenario-based Kubernetes interview runbooks.
Browse All Kubernetes Questions →

📚 Related Production Scenarios in Kubernetes