⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AWS & Cloud Architecture Interview Questions Scenario 145 of 186 in AWS & Cloud Architecture
Senior DevOps / SRE GCP & Cloud Cloud Networking & Triage Incident Postmortem

Q: During a massive Black Friday autoscaling surge, GKE microservices attempting to call external payment gateways (Stripe, PayPal) suddenly experience 30-second connection timeouts. Pods and nodes are healthy. How do you diagnose and fix Cloud NAT port exhaustion?

Root cause analysis and mitigation runbook for SNAT port exhaustion on Google Cloud NAT gateways, preventing outbound packet drops and third-party API timeout cascades during autoscaling surges.

#GCP #Cloud NAT #Networking #GKE #SNAT #Troubleshooting
🎙️ Candidate Opening & Architectural Context
"At 11:30 PM, our e-commerce platform experienced sudden checkout failure spikes. GKE nodes had scaled from 20 to 120 nodes, but outbound TCP handshakes to third-party endpoints were silently timing out."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Examine Cloud NAT Metric Telemetry in Cloud Monitoring

Check SNAT port allocation and dropped connection counts:

  • Metric Check: Inspected compute.googleapis.com/nat/dropped_sent_packets_count with metric filter reason=OUT_OF_RESOURCES.
  • Port Allocation: Monitored compute.googleapis.com/nat/allocated_ports versus open_connections per VM node.
  • Incident Confirmation: Confirmed that newly autoscaled GKE nodes were allocated zero SNAT ports because the Cloud NAT gateway had exhausted its pool of public IP addresses.
Pro Tip: By default, Cloud NAT reserves a fixed number of ports (e.g. 64 ports) per VM. When nodes autoscale rapidly, the public IP pool runs out of 64k port blocks.
2️⃣

Emergency Mitigation: Add Public IPs & Enable Dynamic Port Allocation

Instantly expand the SNAT port capacity pool:

  • Add External IPs: Added 4 reserved static external IP addresses to the Cloud NAT gateway: gcloud compute routers nats update prod-nat --router=prod-router --region=us-central1 --nat-external-ip-pool=nat-ip-1,nat-ip-2,nat-ip-3,nat-ip-4.
  • Enable Dynamic Port Allocation: Enabled Dynamic Port Allocation to allow high-traffic nodes to request more ports as needed rather than wasting static allocations on idle nodes.
  • Command: gcloud compute routers nats update prod-nat --router=prod-router --region=us-central1 --enable-dynamic-port-allocation.
Pro Tip: Dynamic Port Allocation dynamically assigns ports between min-ports-per-vm (32) and max-ports-per-vm (2048), eliminating port starvation.
3️⃣

Tune TCP TIME_WAIT & Connection Re-use Timeouts

Reclaim dead TCP ports aggressively:

  • TCP Established Timeout: Lowered TCP established timeout from 1200s to 300s.
  • TCP Time-Wait: Lowered --tcp-time-wait-timeout from 120s to 30s to rapidly cycle closed sockets.
  • Application Connection Pooling: Fixed an application bug where HTTP clients were creating a new TCP connection per payment request instead of reusing Keep-Alive pools.
Pro Tip: Lowering TCP TIME_WAIT from 120s to 30s frees up thousands of closed ports for immediate reuse during traffic spikes.
4️⃣

Establish Cloud Monitoring Port Saturation Alerting

Prevent future incidents with predictive threshold alerts:

  • Alert Policy: Created alert when allocated_ports / total_available_ports > 75% for more than 3 minutes.
  • Auto-Allocate IPs: Configured --auto-allocate-nat-external-ips on non-restricted NAT routers to let GCP scale external IPs automatically.
Pro Tip: If your third-party APIs enforce IP whitelisting, you cannot use auto-allocate; you must pre-provision static IP blocks and share them with vendors.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Cloud NAT port exhaustion occurs during autoscaling when static port allocations consume all 64k ports of existing NAT IPs; enabling Dynamic Port Allocation and tuning TCP TIME_WAIT permanently resolves the bottleneck."
⚡ 60-Second Elevator Pitch Talking Points
  • Identified SNAT port exhaustion via dropped_sent_packets_count metric with reason OUT_OF_RESOURCES.
  • Executed emergency mitigation by adding static NAT IPs to the gateway pool.
  • Enabled Dynamic Port Allocation so bursting GKE nodes dynamically acquire ports on demand.
  • Reduced TCP TIME_WAIT timeout from 120s to 30s and fixed application HTTP Keep-Alive pooling.
Advertisement
Want more AWS & Cloud Architecture scenarios?
Explore our complete collection of scenario-based AWS & Cloud Architecture interview runbooks.
Browse All AWS & Cloud Architecture Questions →