Q: Your company has a 10 Gbps Direct Connect fiber line from London to Tokyo. However, a single large file transfer using scp/TCP maxes out at only 150 Mbps, despite the link being 99% idle. Why can't TCP fill the pipe, and how do you fix it?
This is classic behavior in a Long Fat Network (LFN)—high bandwidth, high latency.
🛠️ Production Runbook & Step-by-Step Resolution
Production Solution & Architecture
This is classic behavior in a Long Fat Network (LFN)—high bandwidth, high latency. The speed limit is not the bandwidth; it's the TCP Receive Window. TCP requires an acknowledgment (ACK) for data sent. If the window size is 64KB, the sender can only put 64KB of data "in flight" on the fiber cable before stopping to wait for the ACK from Tokyo. Because the round-trip time (ping) from London to Tokyo is huge (e.g., 250ms), the sender constantly stops and waits, wasting the 10Gbps pipe. To fix: I must tune the OS kernel to enable TCP Window Scaling (sysctl net.ipv4.tcp_window_scaling=1) and massively increase the rmem and wmem buffer sizes so the sender can put gigabytes of data "in flight" without waiting for instant ACKs. Changing the congestion control algorithm to BBR also drastically improves throughput on long links.
- Immediate Triage: This is classic behavior in a Long Fat Network (LFN)—high bandwidth, high latency.
- Run targeted verification commands before modifying configuration.
- Automate permanent guardrails (CI check, alerts, IaC policy) to prevent recurrence.