⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 16 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Model Serving & Inference LLM Serving
🎯 Target Role / Context: Staff AI Platform Engineer sizing compute topology and inference clusters for multi-billion parameter foundation models.

Q: When serving a 70B or 405B parameter LLM, how do you decide the split between Tensor Parallelism (TP) and Pipeline Parallelism (PP)? Why should Tensor Parallelism strictly never cross non-NVLink node boundaries, and how does interconnect bandwidth impact token generation latency?

Architectural selection between Tensor Parallelism (TP) and Pipeline Parallelism (PP) in vLLM/SGLang, evaluating NVLink vs PCIe bandwidth constraints and intra-node vs inter-node serving limits.

#Tensor Parallelism #Pipeline Parallelism #vLLM #NVLink #Interconnects #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Modern open-weight models (such as Llama 3 70B, Qwen 72B, DeepSeek-V3) are too large to fit into a single GPU's VRAM. Distributing model weights across multiple GPUs requires model parallelism. However, the two primary forms of model parallelism—Tensor Parallelism (TP) and Pipeline Parallelism (PP)—exhibit radically different network bandwidth and latency sensitivities."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Understand Tensor Parallelism (TP) and Communication Intensity

Tensor Parallelism splits individual weight matrices (Linear, Attention, MLP layers) across GPUs (Megatron-LM style). In every single transformer layer, GPUs must perform two All-Reduce collectives (one after multi-head attention, one after MLP). For a 80-layer model, generating a single token requires 160 All-Reduce operations! Consequently, TP requires ultra-high bandwidth and microsecond latency: it is strictly confined to GPUs connected via NVLink / NVSwitch (900 GB/s on H100).

# vLLM launch with TP=8 (confined to a single 8x GPU node)
python3 -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Meta-Llama-3-70B-Instruct \
  --tensor-parallel-size 8 \
  --pipeline-parallel-size 1
2

Understand Pipeline Parallelism (PP) and Inter-Node Boundary Spanning

Pipeline Parallelism partitions sequential layers of the model across GPUs or nodes (e.g., layers 1-20 on Node 1, layers 21-40 on Node 2). Communication occurs only at layer boundaries: the activations of layer 20 are sent via point-to-point P2P socket communication to Node 2. Because point-to-point transfers require orders of magnitude less bandwidth than 160 All-Reduces per token, PP can safely span across standard network interfaces or RoCE/InfiniBand inter-node boundaries.

# Serving 405B model across two 8x H100 nodes: TP=8 intra-node, PP=2 inter-node
python3 -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Meta-Llama-3-405B-Instruct \
  --tensor-parallel-size 8 \
  --pipeline-parallel-size 2
Advertisement
3

Benchmark Latency Consequences of Crossing Non-NVLink Boundaries with TP

Attempting to run TP across nodes connected over standard 100Gbps Ethernet or PCIe Gen4 introduces catastrophic serialization latency: an All-Reduce that completes in 15 microseconds over NVLink takes 2-5 milliseconds over standard networking. Multiplied by 160 operations per token, generation latency spikes from 30ms/token to over 400ms/token, completely destroying interactive chat SLAs.

# Rule of thumb: TP <= number of GPUs per NVLink domain (usually 8 on HGX H100/A100)
# If Model VRAM > Single Node VRAM: Use TP=8 (intra-node) combined with PP=N (inter-node)
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Tensor Parallelism requires 160 All-Reduce collectives per token and must strictly remain within high-bandwidth NVLink domains. Pipeline Parallelism transmits activations point-to-point across layer boundaries and should be used to span across multi-node boundaries."
⚡ 60-Second Elevator Pitch Talking Points
  • Tensor Parallelism executes two All-Reduce operations per transformer layer for every single generated token.
  • Because of this extreme communication frequency, TP must never cross NVLink boundaries; running TP over PCIe or Ethernet kills token latency by 10x.
  • For ultra-large models like Llama 405B, we configure hybrid parallelism: Tensor Parallelism (TP=8) inside each HGX node via NVLink, and Pipeline Parallelism (PP=2) across nodes over the network.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →