⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 74 of 98 in FinOps & System Design
Staff AI Platform Architect System Design AI Infrastructure & MLOps SRE System Design

Q: Your enterprise is deploying internal Large Language Models (LLaMA 3 70B, Mistral) to serve 5,000 internal developers and customer-facing copilots. Cloud GPU instances (NVIDIA H100/A100) are scarce and cost $4/hr per GPU. How do you design an enterprise LLM serving platform on Kubernetes that maximizes tokens/sec throughput, minimizes Time to First Token (TTFT), and elastically scales GPU clusters based on real-time inference queue depths?

Architectural blueprint for engineering a high-throughput, low-latency LLM inference serving platform on Kubernetes using vLLM PagedAttention, NVIDIA Triton, Ray Serve, and GPU-metric autoscaling.

#System Design #LLM Serving #vLLM #Triton #Kubernetes #GPU Autoscaling #KEDA
🎙️ Candidate Opening & Architectural Context
"Standard container serving frameworks (FastAPI/Flask) fail for LLMs because they do not optimize GPU memory KV caches, resulting in terrible concurrency and massive idle GPU waste. We architected a production LLM serving platform using vLLM, continuous batching, and KEDA queue autoscaling."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Deploy High-Performance Inference Engines with vLLM & PagedAttention

Maximize GPU memory efficiency and throughput using continuous batching:

  • vLLM Inference Engine: Deployed vLLM with PagedAttention, which allocates Key-Value (KV) cache memory in non-contiguous virtual blocks (analogous to virtual memory in operating systems).
  • Continuous Batching: Dynamically adds incoming requests into existing execution batches at iteration level, increasing GPU compute utilization from 25% to 88% and delivering 4x higher tokens/sec throughput.
Pro Tip: PagedAttention virtually eliminates internal GPU memory fragmentation, allowing a single H100 GPU to serve 5x more concurrent requests than standard Hugging Face pipelines.
2️⃣

Configure Tensor Parallelism for Large Models (70B+) via Ray Serve

Shard models larger than a single GPU across multiple NVLink-connected GPUs:

  • Tensor Parallelism: For LLaMA 3 70B (requiring ~140 GB VRAM), configured tensor parallelism tensor_parallel_size: 4 across 4x A100 80GB SXM4 GPUs communicating over high-speed NVLink (600 GB/s).
  • Model Weight Pre-Warming: Stored model checkpoints on local high-performance NVMe shared storage (AWS FSx for Lustre / Google Filestore), slashing pod cold-start weight loading time from 15 minutes down to 35 seconds.
Pro Tip: Tensor parallelism requires high-speed NVLink interconnects between GPUs within the same host to prevent PCIe bus bottlenecks.
3️⃣

Elastic GPU Autoscaling with KEDA & vLLM Queue Telemetry

Scale expensive GPU node pools based on real-time request queue depths rather than CPU/GPU percentage:

  • KEDA Autoscaling: Standard GPU utilization metrics stay at 100% even for a single query. Deployed KEDA scaling on vLLM custom Prometheus metrics: vllm:num_requests_waiting and vllm:avg_prompt_throughput_tok_per_s.
  • Scaling Rules: Spawns an additional GPU pod whenever waiting request queue depth exceeds 10 requests for > 15 seconds.
  • Scale-to-Zero Off-Peak: Scales non-critical fine-tuned models down to 0 replicas at night, saving over $45,000/month in idle GPU reservation costs.
Pro Tip: Autoscaling LLM workloads on waiting request queue depth provides responsive scaling before Time to First Token (TTFT) degrades.
4️⃣

Deploy Semantic Caching (GPTCache) & Guardrail Moderation Layer

Intercept redundant prompts and filter security jailbreaks at the perimeter:

  • Semantic Cache: Deployed GPTCache backed by Redis vector search; queries with >0.96 cosine similarity return cached responses in 15ms with ZERO GPU execution cost (32% cache hit rate).
  • Security Guardrails: Integrated NeMo Guardrails to intercept prompt injections, PII data leakage, and toxic outputs before requests hit downstream models.
Pro Tip: Semantic caching bypasses GPU compute completely for frequent user queries, dramatically reducing inference serving costs.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"An enterprise LLM serving platform pairs vLLM PagedAttention for continuous batching, Ray Serve tensor parallelism over NVLink, KEDA queue-based GPU autoscaling, and semantic caching to maximize tokens/sec under budget."
⚡ 60-Second Elevator Pitch Talking Points
  • Serve models using vLLM PagedAttention to achieve 4x higher throughput via continuous batching.
  • Shard 70B+ parameter models across GPUs using Ray Serve tensor parallelism over NVLink.
  • Autoscale expensive GPU worker pools based on real-time vLLM waiting request queue depths via KEDA.
  • Deploy semantic caching to serve 30% of prompts from cache in 15ms with zero GPU cost.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →