Q: How do FP8, INT8, and INT4 (AWQ vs GPTQ) quantization techniques compare in production serving? Why is FP8 the gold standard on NVIDIA H100 GPUs while AWQ is preferred on A100s, and how does quantization transition LLMs from memory-bandwidth bound to compute bound?
Evaluating and deploying quantized LLMs in production: comparing native FP8 (Ada/Hopper) against INT4/INT8 (AWQ vs GPTQ), analyzing compute vs memory-bound regimes, and measuring perplexity degradation.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Evaluate Hardware Tensor Core Support: A100 (Ampere) vs H100 (Hopper)
NVIDIA Ampere (A100) has native Tensor Core support for FP16, BF16, INT8, and INT4, but lacks native FP8 hardware instructions. In contrast, NVIDIA Hopper (H100) and Ada Lovelace (L40) introduce native 4th-generation Transformer Engine FP8 Tensor Cores, which execute FP8 matrix multiplications at double the throughput (TFLOPS) of FP16 with near-zero perplexity loss.
# vLLM launch using native FP8 on H100
python3 -m vllm.entrypoints.openai.api_server \
--model neuralmagic/Meta-Llama-3-70B-Instruct-FP8 \
--tensor-parallel-size 4 \
--quantization fp8
Compare Weight-Only vs Weight-and-Activation Quantization (AWQ vs SmoothQuant)
Activation outliers make quantizing both weights and activations to INT8 difficult without severe accuracy degradation. AWQ (Activation-aware Weight Quantization) protects the 1% most salient weight channels and quantizes the rest to INT4 (weight-only). SmoothQuant / FP8 quantizes both weights and activations, unlocking both memory bandwidth savings and accelerated compute GEMM kernels.
# vLLM launch using AWQ INT4 on A100
python3 -m vllm.entrypoints.openai.api_server \
--model casperhansen/llama-3-70b-instruct-awq \
--quantization awq \
--tensor-parallel-size 4
Analyze Memory Bandwidth vs Compute Bound Regimes (Roofline Model)
Apply the Roofline model: At low batch sizes (batch = 1-8), generation is strictly memory-bandwidth bound. INT4/AWQ cuts weights from 140GB to 35GB, resulting in near-linear 3x throughput improvements. At high batch sizes (batch = 64+), arithmetic intensity increases and the workload becomes compute-bound. Here, FP8 Tensor Cores dominate because they double raw computational TFLOPS, whereas INT4 dequantization kernels introduce ALU overhead.
# Production Selection Matrix:
# - High concurrency on H100: Native FP8 (Transformer Engine)
# - Low-to-medium concurrency on A100/T4: AWQ INT4
# - Safety-critical / zero-degradation workloads: BF16
- LLM inference at low batch sizes is bounded by memory bandwidth, not compute power.
- On NVIDIA H100, we use native FP8: Hopper's Transformer Engine executes FP8 GEMMs at 2x the speed of FP16 with near-zero accuracy loss.
- On A100 clusters, we use AWQ INT4, shrinking 70B models into 35GB of VRAM and boosting tokens-per-second by 2.8x.