Q: A training pod crashes with 'torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB (GPU 0; 79.15 GiB total capacity; 42.30 GiB already allocated; 31.85 GiB reserved in PyTorch; 4.00 GiB free)'. Why did this OOM occur when 36.85 GiB appears physically available, and how do you resolve it?
Differentiating between physical VRAM exhaustion and virtual memory address fragmentation in PyTorch, and tuning the PyTorch CUDA caching allocator for zero-OOM execution.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Analyze Allocated vs Reserved vs Free GPU Memory
Deconstruct the error message: Total VRAM is 79.15 GiB. Allocated (active tensors) is 42.30 GiB. Reserved (cached by PyTorch allocator) is 74.15 GiB (42.30 + 31.85). Free according to CUDA driver is only 4.00 GiB. PyTorch tried to allocate 512 MiB: it searched its 31.85 GiB cache, found no contiguous 512 MiB free block due to fragmentation, called cudaMalloc for a new block, and failed because the driver only had 4.00 GiB left and system allocations prevented it.
# Print memory summary in Python
import torch
print(torch.cuda.memory_summary(device=0, abbreviated=False))
Activate PyTorch Expandable Segments to Eliminate Fragmentation
In PyTorch 2.1+, set `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. This enables PyTorch to use CUDA Virtual Memory Management (VMM) APIs (`cuMemCreate`, `cuMemMap`). Instead of requiring physically contiguous memory blocks in the caching allocator, PyTorch dynamically maps virtual memory pages, effectively eradicating allocator fragmentation.
# Set environment variable in Dockerfile or Kubernetes Pod spec
env:
- name: PYTORCH_CUDA_ALLOC_CONF
value: 'expandable_segments:True,max_split_size_mb:128'
Implement Activation Checkpointing and Garbage Collection Hooks
For deep transformer layers, enable activation checkpointing (gradient checkpointing) using `torch.utils.checkpoint.checkpoint`. This drops intermediate activations during the forward pass and recomputes them during the backward pass, reducing peak VRAM requirements by 60-70% at the cost of ~20-30% additional compute.
from torch.utils.checkpoint import checkpoint
class TransformerBlock(torch.nn.Module):
def forward(self, x):
# Recomputes activations on backward pass rather than holding in VRAM
return checkpoint(self._forward_impl, x, use_reentrant=False)
- An OOM with gigabytes of free memory is almost always caused by PyTorch caching allocator fragmentation, where no contiguous block fits the request.
- We enable `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which uses CUDA Virtual Memory APIs to map fragmented physical chunks into contiguous virtual addresses.
- Combined with activation checkpointing, this completely eliminates fragmentation OOMs and allows training 30% larger batch sizes on the same hardware.