Q: How do DeepSpeed ZeRO-3 and PyTorch FSDP shard optimizer states, gradients, and model parameters across GPUs, and when should you enable CPU/NVMe memory offloading versus accepting communication bandwidth bottlenecks?
Comparing Zero Redundancy Optimizer stages (ZeRO-1/2/3) against PyTorch Fully Sharded Data Parallel (FSDP), communication overhead trade-offs, and host CPU/NVMe memory offloading.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deconstruct Sharding Stages: ZeRO-1, ZeRO-2, ZeRO-3 (FSDP FULL_SHARD)
ZeRO-1 shards optimizer states (saving 4x memory with Adam). ZeRO-2 shards optimizer states and gradients (saving up to 8x memory). ZeRO-3 (and PyTorch FSDP in FULL_SHARD mode) shards optimizer states, gradients, and model parameters: each GPU retains only 1/N of the model weights. During the forward pass, an All-Gather collective fetches the missing parameters for layer L, computes activations, and immediately discards the gathered weights to conserve VRAM.
# PyTorch FSDP sharding strategy definition
from torch.distributed.fsdp import FullyShardedDataParallel as FSDP, ShardingStrategy
fsdp_model = FSDP(
model,
sharding_strategy=ShardingStrategy.FULL_SHARD, # ZeRO-3 equivalent
auto_wrap_policy=transformer_auto_wrap_policy
)
Evaluate Communication Volume vs Memory Trade-Off
ZeRO-3/FSDP increases total network communication by 50% compared to standard DDP: DDP performs a single All-Reduce on gradients during the backward pass (2x model size), whereas ZeRO-3 requires an All-Gather in the forward pass, an All-Gather in the backward pass, and a Reduce-Scatter on gradients (totaling 3x model size). Without high-speed interconnects (NVLink or 400Gbps RoCE), ZeRO-3 can become severely network-bound.
# NCCL communication volume rule:
# DDP: 2 * Parameters (AllReduce)
# ZeRO-3 / FSDP: 3 * Parameters (AllGather forward + AllGather backward + ReduceScatter)
Configure CPU and NVMe Parameter Offloading
When fine-tuning giant models that exceed cluster aggregate VRAM, configure DeepSpeed ZeRO-Offload or FSDP CPU offload. Optimizer states and master weights are placed in host CPU RAM or stripped NVMe drives. Only the active layer weights are streamed over PCIe to the GPU. This allows fine-tuning 70B models on a single 8x A100 node at the expense of PCIe transfer latency.
# DeepSpeed ds_config.json offload configuration
"zero_optimization": {
"stage": 3,
"offload_optimizer": {
"device": "cpu",
"pin_memory": true
},
"offload_param": {
"device": "cpu",
"pin_memory": true
}
}
- DDP clones the whole model on every GPU, while ZeRO-3 and PyTorch FSDP shard weights, gradients, and optimizer states across ranks.
- This cuts memory consumption by N times, requiring All-Gather collectives right before each layer executes.
- When aggregate VRAM is insufficient, we enable CPU offloading over host pinned RAM, trading PCIe transfer latency for the ability to train 70B models on smaller clusters.