Q: A request is going from Pod A to Pod B via a Service and it's very slow. How do you troubleshoot network latency in Kubernetes?
1. Baseline test — kubectl exec -it <pod-a> -- curl -o /dev/null -s -w "%{time_total}" http://<service>:<port> to measure actual latency.
#Kubernetes #Networking #L3 #Container Orchestration #K8s
🎙️ Candidate Opening & Architectural Context
""When troubleshooting Kubernetes, I always follow a structured layered model: Pod status -> Events -> Logs -> Network. When addressing this question, I walk the interviewer through our production incident runbook: isolating the blast radius, checking diagnostic logs and metrics, and applying a safe fix.""
Advertisement
🛠️ Production Runbook & Step-by-Step Resolution
1️⃣
Initial Diagnostics & Root Cause Analysis
- Baseline test —
kubectl exec -itto measure actual latency.-- curl -o /dev/null -s -w "%{time_total}" http:// : - Bypass the service — test pod-to-pod directly using the pod IP to see if latency is in kube-proxy/iptables:
kubectl exec -it.-- curl http:// : - Check kube-proxy mode — iptables vs ipvs. ipvs is faster at scale.
- Check CNI — network plugin issues. Run
pingbetween pods to test raw network latency.
2️⃣
Remediation & Permanent Safeguards
Execute the resolution runbook and verify workload health:
- DNS latency —
kubectl exec -it— DNS lookups through CoreDNS add latency. Consider-- time nslookup ndots:5setting impact. - Node-level network — check if nodes are on the same AZ. Cross-AZ traffic adds ~1-2ms.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Baseline test — kubectl exec -it -- curl -o /dev/null -s -w "%{time_total}" http://: to measure actual latency.."
⚡ 60-Second Elevator Pitch Talking Points
- Baseline test — kubectl exec -it -- curl -o /dev/null -s -w "%{time_total}" http://: to measure ...
- Bypass the service — test pod-to-pod directly using the pod IP to see if latency is in kube-proxy...
- Check kube-proxy mode — iptables vs ipvs. ipvs is faster at scale.
Advertisement