Q: A fleet of 5,000 Lambda functions in a private VPC aggressively scrape data from the public internet. Randomly, hundreds of them begin crashing with bizarre `Connection Timed Out` networking errors, despite the internet destination being perfectly healthy. What AWS bottleneck is occurring?
This is classic SNAT (Source Network Address Translation) Port Exhaustion on the NAT Gateway.
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
This is classic SNAT (Source Network Address Translation) Port Exhaustion on the NAT Gateway. A single AWS NAT Gateway utilizes a single public Elastic IP. TCP allows a theoretical maximum of ~65,000 ephemeral outbound ports per IP addressing a single destination. When 5,000 highly concurrent Lambda functions open thousands of individual API connections to the exact same external internet API simultaneously, the NAT Gateway completely runs out of ephemeral routing ports. It violently drops any new outbound connection attempts until old ones close. *Fix:* Heavily deploy multiple NAT Gateways across multiple public subnets and route traffic dynamically to distribute the SNAT allocation, or deploy dedicated NAT instances.
- Immediate Triage: This is classic SNAT (Source Network Address Translation) Port Exhaustion on the NAT Gateway.
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.