Q: Why does colocating prompt prefill and token generation on the same GPUs cause catastrophic tail latency and poor hardware utilization? How do disaggregated serving architectures (like SplitWise / Mooncake / vLLM Disaggregated) decouple prefill from decode over RDMA?
Deep architectural analysis of disaggregated LLM serving: separating compute-bound prompt prefill nodes from memory-bandwidth-bound token decode nodes over high-speed RDMA KV cache transfer fabrics.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Identify the Prefill vs Decode Interference Bottleneck
When a long 4,000-token prompt arrives at an active GPU, the prefill computation occupies the Tensor Cores for 300-500 milliseconds. Active streaming users experiencing real-time token decoding are starved of compute cycles: their Inter-Token Latency (ITL) spikes from 25ms to 500ms, causing noticeable stuttering in interactive chat interfaces.
# Co-located serving problem:
# Active Decoding Request: Token 1 (25ms) -> Token 2 (25ms) -> [LONG PREFILL ARRIVES: 450ms freeze] -> Token 3...
Architect Disaggregated Compute Clusters (Prefill Pool vs Decode Pool)
Partition infrastructure into two distinct GPU pools: (1) Prefill Cluster: Sized with compute-optimized GPUs (e.g. NVIDIA H100 with high TFLOPS), running large batch sizes to process prompts at maximum arithmetic saturation. (2) Decode Cluster: Sized with memory-capacity and bandwidth-optimized GPUs (e.g. H200 / L40S with massive HBM/VRAM), dedicated exclusively to low-latency token generation without interruption.
# Architecture Flow:
# Client -> API Gateway -> Prefill Worker Pool (Computes KV Cache)
# Prefill Worker -> RDMA Transfer (Transfers KV Cache) -> Decode Worker Pool
# Decode Worker Pool -> Streams tokens back to client
High-Speed RDMA KV Cache Transfer Protocol
The linchpin of disaggregated serving is KV cache transfer latency. Once the Prefill Worker finishes prompt computation, it must transfer the generated KV cache tensors to the assigned Decode Worker. Using standard TCP/HTTP incurs tens of milliseconds of serialization delay. Modern frameworks (vLLM Disaggregated, Mooncake) implement zero-copy RDMA or NVLink over InfiniBand/RoCE, transferring gigabytes of KV cache in under 2 milliseconds.
# vLLM disaggregated prefill-decode prototype configuration
# Prefill worker launches with: --is-prefill-worker
# Decode worker launches with: --is-decode-worker --kv-transfer-protocol rdma
- Processing a long prompt on a shared GPU freezes all active streaming users for half a second.
- We disaggregate our serving infrastructure: a dedicated Prefill cluster crunches prompts at peak Tensor Core saturation, while a Decode cluster streams tokens with zero interruption.
- By streaming KV caches between clusters over zero-copy RDMA in under 2 milliseconds, we cut p99 inter-token latency spikes by 80% while doubling overall GPU utilization.