Q: Your Kubernetes cluster is experiencing intermittent pod failures and high latency. How would you troubleshoot it systematically? Explain how you would investigate pods, nodes, networking, resource limits, probes, scheduling, DNS, and application metrics.
Exhaustive diagnostic methodology for intermittent Kubernetes failures: node kernel pressure, CPU throttling, CoreDNS latency (ndots:5), CNI IP exhaustion, and conntrack table drops.
#Kubernetes #CoreDNS #Latency #Troubleshooting #Conntrack #Throttling #CNI
🎙️ Candidate Opening & Architectural Context
"Intermittent failures and latency spikes in Kubernetes are notoriously elusive because they rarely show up as hard crashes. I isolate them using a layered, full-stack diagnostic model: Nodes -> Pods & Limits -> CoreDNS -> Networking/CNI -> Application Probes."
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Layer 1: Node Health, Kernel & CPU Throttling
Inspect underlying host instances and container runtime:
- Run
kubectl get nodes: Check for node conditions likeMemoryPressure,DiskPressure, orPIDPressure. - SSH into suspected nodes: check
dmesg -Tfor kernel OOM-killer invocations, hardware errors, or TCP drops. - CPU Throttling Trap: Check Prometheus metric
container_cpu_cfs_throttled_seconds_total. Even when average node CPU is only 40%, strict podresources.limits.cpucause Linux CFS throttling, introducing random 200–500ms latency spikes!
2️⃣
Layer 2: Pod Restarts, OOMKills & Exit Codes
Inspect container states across namespaces:
kubectl get pods -A --sort-by='.status.containerStatuses[0].restartCount': Identify flapping pods.kubectl describe pod <pod>: Check forOOMKilled(Exit Code 137).- Retrieve stack trace from crashed container:
kubectl logs <pod> -c <container> --previous. - Check liveness probe thresholds: Are slow database calls causing the liveness probe to timeout and restart otherwise healthy pods?
3️⃣
Layer 3: CoreDNS & The 'ndots:5' DNS Latency Trap
Intermittent 1-second latency spikes are almost always DNS issues:
- Check CoreDNS latency:
coredns_dns_request_duration_secondsin Prometheus. - The ndots:5 Amplification Trap: Default Kubernetes
/etc/resolv.confhasndots:5. For external domains (e.g.api.stripe.com), the pod queries 4 internal search domains (.default.svc...) before querying the public domain, multiplying DNS queries by 5x and overloading CoreDNS! - Fix: Deploy NodeLocal DNSCache daemonset on all nodes, or append a trailing dot (
api.stripe.com.) in application configs.
4️⃣
Layer 4: CNI, IP Exhaustion & Conntrack Table
Subtle networking drops at the host and CNI level:
- VPC CNI IP Exhaustion: In AWS, check if worker node subnets ran out of free private IP addresses, preventing newly scheduled pods from obtaining an ENI secondary IP.
- Linux Conntrack Saturation: High-traffic microservices exhaust the Linux connection tracking table (
nf_conntrack_max). Once full, the kernel silently drops new TCP SYN packets! Check withdmesg -T | grep 'table full, dropping packet'. - kube-proxy Sync Latency: Check if iptables rule processing is stalling packet forwarding.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Intermittent K8s latency is usually not pod crashes: it's Linux CFS CPU throttling, CoreDNS ndots:5 search domain multiplication, or Linux nf_conntrack table exhaustion. Deploy NodeLocal DNSCache and tune CPU limits."
⚡ 60-Second Elevator Pitch Talking Points
- Layer 1 (Node): Check node pressure flags (Memory/DiskPressure) and CFS CPU throttling (container_cpu_cfs_throttled_seconds_total).
- Layer 2 (Pods): Sort pods by restart count; inspect 'kubectl logs --previous' for OOMKilled (Exit Code 137).
- Layer 3 (DNS): Inspect CoreDNS metrics; mitigate ndots:5 search domain amplification using NodeLocal DNSCache.
- Layer 4 (Network): Verify VPC CNI subnet IP availability; check 'dmesg' for nf_conntrack table exhaustion drops.
- Layer 5 (Probes): Ensure liveness probes have sufficient timeout/initialDelay to avoid killing slow-starting pods.
Advertisement