⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 50 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Model Serving & Inference LLM Serving
🎯 Target Role / Context: Staff AI Infrastructure Architect designing next-generation hyperscale foundation model serving platforms.

Q: Why does colocating prompt prefill and token generation on the same GPUs cause catastrophic tail latency and poor hardware utilization? How do disaggregated serving architectures (like SplitWise / Mooncake / vLLM Disaggregated) decouple prefill from decode over RDMA?

Deep architectural analysis of disaggregated LLM serving: separating compute-bound prompt prefill nodes from memory-bandwidth-bound token decode nodes over high-speed RDMA KV cache transfer fabrics.

#Disaggregated Serving #SplitWise #Mooncake #Prefill-Decode #vLLM #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"In traditional LLM serving engines, a single GPU or tensor-parallel group handles both the prompt prefill phase (processing input tokens) and the token decode phase (generating new tokens autoregressively). However, prefill and decode have diametrically opposing computational characteristics: prefill is compute-bound (saturating Tensor Cores with high arithmetic intensity), while decode is memory-bandwidth bound (saturating HBM streaming). Conflating them causes severe latency spikes and poor hardware utilization."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Identify the Prefill vs Decode Interference Bottleneck

When a long 4,000-token prompt arrives at an active GPU, the prefill computation occupies the Tensor Cores for 300-500 milliseconds. Active streaming users experiencing real-time token decoding are starved of compute cycles: their Inter-Token Latency (ITL) spikes from 25ms to 500ms, causing noticeable stuttering in interactive chat interfaces.

# Co-located serving problem:
# Active Decoding Request: Token 1 (25ms) -> Token 2 (25ms) -> [LONG PREFILL ARRIVES: 450ms freeze] -> Token 3...
2

Architect Disaggregated Compute Clusters (Prefill Pool vs Decode Pool)

Partition infrastructure into two distinct GPU pools: (1) Prefill Cluster: Sized with compute-optimized GPUs (e.g. NVIDIA H100 with high TFLOPS), running large batch sizes to process prompts at maximum arithmetic saturation. (2) Decode Cluster: Sized with memory-capacity and bandwidth-optimized GPUs (e.g. H200 / L40S with massive HBM/VRAM), dedicated exclusively to low-latency token generation without interruption.

# Architecture Flow:
# Client -> API Gateway -> Prefill Worker Pool (Computes KV Cache)
# Prefill Worker -> RDMA Transfer (Transfers KV Cache) -> Decode Worker Pool
# Decode Worker Pool -> Streams tokens back to client
Advertisement
3

High-Speed RDMA KV Cache Transfer Protocol

The linchpin of disaggregated serving is KV cache transfer latency. Once the Prefill Worker finishes prompt computation, it must transfer the generated KV cache tensors to the assigned Decode Worker. Using standard TCP/HTTP incurs tens of milliseconds of serialization delay. Modern frameworks (vLLM Disaggregated, Mooncake) implement zero-copy RDMA or NVLink over InfiniBand/RoCE, transferring gigabytes of KV cache in under 2 milliseconds.

# vLLM disaggregated prefill-decode prototype configuration
# Prefill worker launches with: --is-prefill-worker
# Decode worker launches with: --is-decode-worker --kv-transfer-protocol rdma
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Disaggregated serving separates compute-bound prompt prefill from memory-bandwidth-bound token decode onto specialized GPU pools. Transferring KV caches over low-latency RDMA eliminates decode stutter and increases cluster throughput by up to 2x."
⚡ 60-Second Elevator Pitch Talking Points
  • Processing a long prompt on a shared GPU freezes all active streaming users for half a second.
  • We disaggregate our serving infrastructure: a dedicated Prefill cluster crunches prompts at peak Tensor Core saturation, while a Decode cluster streams tokens with zero interruption.
  • By streaming KV caches between clusters over zero-copy RDMA in under 2 milliseconds, we cut p99 inter-token latency spikes by 80% while doubling overall GPU utilization.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →