Q: You see a gradual increase in latency across an NGINX ingress controller, but CPU and memory are completely stable. What’s your next move?
Deep dive into NGINX Ingress controller latency regressions occurring under stable CPU and memory utilization due to socket backlog saturation and upstream connection churn.
Want to master this scenario in a live sandbox? KodeKloud's Istio Service Mesh & Advanced Kubernetes Networking Course covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Check Upstream Keepalive Configuration (Connection Churn)
By default, NGINX Ingress establishes a brand new TCP connection (and TLS handshake) for EVERY request sent to backend pods unless `upstream-keepalive-connections` is explicitly enabled. Under high RPS, this triggers massive ephemeral port exhaustion and TCP handshake latency spikes while CPU remains low.
# Enable upstream keepalive in Ingress ConfigMap
apiVersion: v1
kind: ConfigMap
metadata:
name: ingress-nginx-controller
namespace: ingress-nginx
data:
upstream-keepalive-connections: "100"
upstream-keepalive-timeout: "60"
upstream-keepalive-requests: "10000"
Inspect Socket Listen Backlog & Drops (netstat -s / ss -lnt)
Check if the operating system is dropping incoming TCP SYN packets because the socket listen backlog is saturated.
kubectl exec -it <nginx-ingress-pod> -- netstat -s | grep -i "listen"
# Look for: "times the listen queue of a socket overflowed" or "SYNs to LISTEN sockets dropped"
Inspect NGINX Worker Connections & Event Loop Limits
Check whether NGINX is hitting `worker-connections` (default 16384). If saturated, new client requests queue up in OS buffers, causing perceived latency to soar.
kubectl logs <nginx-ingress-pod> -n ingress-nginx | grep -i "worker_connections"
Verify Dynamic Upstream Service Endpoint Churn
If backend application pods are frequently scaling or restarting, NGINX Ingress constantly reloads or queries the kube-apiserver endpoints controller. Enable Lua-based dynamic upstream routing to eliminate configuration reload latency.
- Inspect upstream keepalives: NGINX defaults to opening a fresh TCP connection per request, choking on handshakes.
- Check kernel socket listen queue overflows via netstat -s in the ingress pod.
- Verify NGINX worker-connections limits and tune net.core.somaxconn.
- Ensure dynamic endpoint routing is enabled to avoid frequent configuration reloads during autoscaling.