⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 45 of 50 in AI/ML Infrastructure & GPU
Staff AI Infrastructure Engineer AI/ML Infrastructure Model Serving & Quantization Model Quantization
🎯 Target Role / Context: Staff AI Infrastructure Engineer optimizing inference efficiency, VRAM capacity, and serving costs across diverse GPU generations.

Q: How do FP8, INT8, and INT4 (AWQ vs GPTQ) quantization techniques compare in production serving? Why is FP8 the gold standard on NVIDIA H100 GPUs while AWQ is preferred on A100s, and how does quantization transition LLMs from memory-bandwidth bound to compute bound?

Evaluating and deploying quantized LLMs in production: comparing native FP8 (Ada/Hopper) against INT4/INT8 (AWQ vs GPTQ), analyzing compute vs memory-bound regimes, and measuring perplexity degradation.

#Quantization #FP8 #AWQ #GPTQ #Tensor Cores #vLLM #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Generative LLM decoding is predominantly memory-bandwidth bound: for batch sizes under 32, the GPU spends most of its time streaming weight matrices from high-bandwidth memory (HBM) to on-chip SRAM rather than computing math. Quantization compresses weights from 16-bit (FP16/BF16) down to 8-bit (FP8/INT8) or 4-bit (INT4), cutting memory traffic by 2x to 4x and dramatically increasing throughput."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Evaluate Hardware Tensor Core Support: A100 (Ampere) vs H100 (Hopper)

NVIDIA Ampere (A100) has native Tensor Core support for FP16, BF16, INT8, and INT4, but lacks native FP8 hardware instructions. In contrast, NVIDIA Hopper (H100) and Ada Lovelace (L40) introduce native 4th-generation Transformer Engine FP8 Tensor Cores, which execute FP8 matrix multiplications at double the throughput (TFLOPS) of FP16 with near-zero perplexity loss.

# vLLM launch using native FP8 on H100
python3 -m vllm.entrypoints.openai.api_server \
  --model neuralmagic/Meta-Llama-3-70B-Instruct-FP8 \
  --tensor-parallel-size 4 \
  --quantization fp8
2

Compare Weight-Only vs Weight-and-Activation Quantization (AWQ vs SmoothQuant)

Activation outliers make quantizing both weights and activations to INT8 difficult without severe accuracy degradation. AWQ (Activation-aware Weight Quantization) protects the 1% most salient weight channels and quantizes the rest to INT4 (weight-only). SmoothQuant / FP8 quantizes both weights and activations, unlocking both memory bandwidth savings and accelerated compute GEMM kernels.

# vLLM launch using AWQ INT4 on A100
python3 -m vllm.entrypoints.openai.api_server \
  --model casperhansen/llama-3-70b-instruct-awq \
  --quantization awq \
  --tensor-parallel-size 4
Advertisement
3

Analyze Memory Bandwidth vs Compute Bound Regimes (Roofline Model)

Apply the Roofline model: At low batch sizes (batch = 1-8), generation is strictly memory-bandwidth bound. INT4/AWQ cuts weights from 140GB to 35GB, resulting in near-linear 3x throughput improvements. At high batch sizes (batch = 64+), arithmetic intensity increases and the workload becomes compute-bound. Here, FP8 Tensor Cores dominate because they double raw computational TFLOPS, whereas INT4 dequantization kernels introduce ALU overhead.

# Production Selection Matrix:
# - High concurrency on H100: Native FP8 (Transformer Engine)
# - Low-to-medium concurrency on A100/T4: AWQ INT4
# - Safety-critical / zero-degradation workloads: BF16
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"On NVIDIA H100, native FP8 quantization is optimal because Hopper Tensor Cores double compute TFLOPS with minimal perplexity loss. On A100, AWQ INT4 is preferred to slash memory bandwidth consumption during memory-bound decoding."
⚡ 60-Second Elevator Pitch Talking Points
  • LLM inference at low batch sizes is bounded by memory bandwidth, not compute power.
  • On NVIDIA H100, we use native FP8: Hopper's Transformer Engine executes FP8 GEMMs at 2x the speed of FP16 with near-zero accuracy loss.
  • On A100 clusters, we use AWQ INT4, shrinking 70B models into 35GB of VRAM and boosting tokens-per-second by 2.8x.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →