Q: When serving a 70B or 405B parameter LLM, how do you decide the split between Tensor Parallelism (TP) and Pipeline Parallelism (PP)? Why should Tensor Parallelism strictly never cross non-NVLink node boundaries, and how does interconnect bandwidth impact token generation latency?
Architectural selection between Tensor Parallelism (TP) and Pipeline Parallelism (PP) in vLLM/SGLang, evaluating NVLink vs PCIe bandwidth constraints and intra-node vs inter-node serving limits.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Understand Tensor Parallelism (TP) and Communication Intensity
Tensor Parallelism splits individual weight matrices (Linear, Attention, MLP layers) across GPUs (Megatron-LM style). In every single transformer layer, GPUs must perform two All-Reduce collectives (one after multi-head attention, one after MLP). For a 80-layer model, generating a single token requires 160 All-Reduce operations! Consequently, TP requires ultra-high bandwidth and microsecond latency: it is strictly confined to GPUs connected via NVLink / NVSwitch (900 GB/s on H100).
# vLLM launch with TP=8 (confined to a single 8x GPU node)
python3 -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-70B-Instruct \
--tensor-parallel-size 8 \
--pipeline-parallel-size 1
Understand Pipeline Parallelism (PP) and Inter-Node Boundary Spanning
Pipeline Parallelism partitions sequential layers of the model across GPUs or nodes (e.g., layers 1-20 on Node 1, layers 21-40 on Node 2). Communication occurs only at layer boundaries: the activations of layer 20 are sent via point-to-point P2P socket communication to Node 2. Because point-to-point transfers require orders of magnitude less bandwidth than 160 All-Reduces per token, PP can safely span across standard network interfaces or RoCE/InfiniBand inter-node boundaries.
# Serving 405B model across two 8x H100 nodes: TP=8 intra-node, PP=2 inter-node
python3 -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-405B-Instruct \
--tensor-parallel-size 8 \
--pipeline-parallel-size 2
Benchmark Latency Consequences of Crossing Non-NVLink Boundaries with TP
Attempting to run TP across nodes connected over standard 100Gbps Ethernet or PCIe Gen4 introduces catastrophic serialization latency: an All-Reduce that completes in 15 microseconds over NVLink takes 2-5 milliseconds over standard networking. Multiplied by 160 operations per token, generation latency spikes from 30ms/token to over 400ms/token, completely destroying interactive chat SLAs.
# Rule of thumb: TP <= number of GPUs per NVLink domain (usually 8 on HGX H100/A100)
# If Model VRAM > Single Node VRAM: Use TP=8 (intra-node) combined with PP=N (inter-node)
- Tensor Parallelism executes two All-Reduce operations per transformer layer for every single generated token.
- Because of this extreme communication frequency, TP must never cross NVLink boundaries; running TP over PCIe or Ethernet kills token latency by 10x.
- For ultra-large models like Llama 405B, we configure hybrid parallelism: Tensor Parallelism (TP=8) inside each HGX node via NVLink, and Pipeline Parallelism (PP=2) across nodes over the network.