⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Networking & Cloud DNS Interview Questions Scenario 52 of 53 in Networking & Cloud DNS
Senior SRE / Networking Specialist Networking Ingress Controllers & Reverse Proxies Tesla Scale Loop

Q: You see a gradual increase in latency across an NGINX ingress controller, but CPU and memory are completely stable. What’s your next move?

Deep dive into NGINX Ingress controller latency regressions occurring under stable CPU and memory utilization due to socket backlog saturation and upstream connection churn.

#Networking #NGINX #Ingress #Keepalive #Worker Connections #Tesla #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"When NGINX Ingress latency climbs steadily while CPU and memory remain flat, NGINX is not bottlenecking on computation. It is bottlenecking on socket connection management: upstream TCP connection thrashing (lack of upstream keepalives), worker connection pool exhaustion, Linux kernel socket listen backlog overflow (`somaxconn`), or DNS resolution delays on upstream service endpoints."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's Istio Service Mesh & Advanced Kubernetes Networking Course covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Check Upstream Keepalive Configuration (Connection Churn)

By default, NGINX Ingress establishes a brand new TCP connection (and TLS handshake) for EVERY request sent to backend pods unless `upstream-keepalive-connections` is explicitly enabled. Under high RPS, this triggers massive ephemeral port exhaustion and TCP handshake latency spikes while CPU remains low.

# Enable upstream keepalive in Ingress ConfigMap
apiVersion: v1
kind: ConfigMap
metadata:
  name: ingress-nginx-controller
  namespace: ingress-nginx
data:
  upstream-keepalive-connections: "100"
  upstream-keepalive-timeout: "60"
  upstream-keepalive-requests: "10000"
2

Inspect Socket Listen Backlog & Drops (netstat -s / ss -lnt)

Check if the operating system is dropping incoming TCP SYN packets because the socket listen backlog is saturated.

kubectl exec -it <nginx-ingress-pod> -- netstat -s | grep -i "listen"
# Look for: "times the listen queue of a socket overflowed" or "SYNs to LISTEN sockets dropped"
Advertisement
3

Inspect NGINX Worker Connections & Event Loop Limits

Check whether NGINX is hitting `worker-connections` (default 16384). If saturated, new client requests queue up in OS buffers, causing perceived latency to soar.

kubectl logs <nginx-ingress-pod> -n ingress-nginx | grep -i "worker_connections"
4

Verify Dynamic Upstream Service Endpoint Churn

If backend application pods are frequently scaling or restarting, NGINX Ingress constantly reloads or queries the kube-apiserver endpoints controller. Enable Lua-based dynamic upstream routing to eliminate configuration reload latency.

Pro Tip: Golden Configuration: Setting upstream-keepalive-connections cuts NGINX Ingress latency by up to 70% during traffic surges.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Ingress latency with flat CPU points to TCP connection thrashing. Enable upstream-keepalive-connections, increase max worker connections, and tune kernel net.core.somaxconn backlog queues."
⚡ 60-Second Elevator Pitch Talking Points
  • Inspect upstream keepalives: NGINX defaults to opening a fresh TCP connection per request, choking on handshakes.
  • Check kernel socket listen queue overflows via netstat -s in the ingress pod.
  • Verify NGINX worker-connections limits and tune net.core.somaxconn.
  • Ensure dynamic endpoint routing is enabled to avoid frequent configuration reloads during autoscaling.
Advertisement
Want more Networking & Cloud DNS scenarios?
Explore our complete collection of scenario-based Networking & Cloud DNS interview runbooks.
Browse All Networking & Cloud DNS Questions →