Q: During a massive Black Friday autoscaling surge, GKE microservices attempting to call external payment gateways (Stripe, PayPal) suddenly experience 30-second connection timeouts. Pods and nodes are healthy. How do you diagnose and fix Cloud NAT port exhaustion?
Root cause analysis and mitigation runbook for SNAT port exhaustion on Google Cloud NAT gateways, preventing outbound packet drops and third-party API timeout cascades during autoscaling surges.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Examine Cloud NAT Metric Telemetry in Cloud Monitoring
Check SNAT port allocation and dropped connection counts:
- Metric Check: Inspected
compute.googleapis.com/nat/dropped_sent_packets_countwith metric filterreason=OUT_OF_RESOURCES. - Port Allocation: Monitored
compute.googleapis.com/nat/allocated_portsversusopen_connectionsper VM node. - Incident Confirmation: Confirmed that newly autoscaled GKE nodes were allocated zero SNAT ports because the Cloud NAT gateway had exhausted its pool of public IP addresses.
Emergency Mitigation: Add Public IPs & Enable Dynamic Port Allocation
Instantly expand the SNAT port capacity pool:
- Add External IPs: Added 4 reserved static external IP addresses to the Cloud NAT gateway:
gcloud compute routers nats update prod-nat --router=prod-router --region=us-central1 --nat-external-ip-pool=nat-ip-1,nat-ip-2,nat-ip-3,nat-ip-4. - Enable Dynamic Port Allocation: Enabled Dynamic Port Allocation to allow high-traffic nodes to request more ports as needed rather than wasting static allocations on idle nodes.
- Command:
gcloud compute routers nats update prod-nat --router=prod-router --region=us-central1 --enable-dynamic-port-allocation.
Tune TCP TIME_WAIT & Connection Re-use Timeouts
Reclaim dead TCP ports aggressively:
- TCP Established Timeout: Lowered TCP established timeout from 1200s to 300s.
- TCP Time-Wait: Lowered
--tcp-time-wait-timeoutfrom 120s to 30s to rapidly cycle closed sockets. - Application Connection Pooling: Fixed an application bug where HTTP clients were creating a new TCP connection per payment request instead of reusing Keep-Alive pools.
Establish Cloud Monitoring Port Saturation Alerting
Prevent future incidents with predictive threshold alerts:
- Alert Policy: Created alert when
allocated_ports / total_available_ports > 75%for more than 3 minutes. - Auto-Allocate IPs: Configured
--auto-allocate-nat-external-ipson non-restricted NAT routers to let GCP scale external IPs automatically.
- Identified SNAT port exhaustion via dropped_sent_packets_count metric with reason OUT_OF_RESOURCES.
- Executed emergency mitigation by adding static NAT IPs to the gateway pool.
- Enabled Dynamic Port Allocation so bursting GKE nodes dynamically acquire ports on demand.
- Reduced TCP TIME_WAIT timeout from 120s to 30s and fixed application HTTP Keep-Alive pooling.