⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 3 of 50 in AI/ML Infrastructure & GPU
Senior AI Infrastructure Engineer AI/ML Infrastructure GPU Architecture & Memory GPU Memory
🎯 Target Role / Context: Senior AI Infrastructure Engineer optimizing deep learning workload reliability and memory utilization.

Q: A training pod crashes with 'torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB (GPU 0; 79.15 GiB total capacity; 42.30 GiB already allocated; 31.85 GiB reserved in PyTorch; 4.00 GiB free)'. Why did this OOM occur when 36.85 GiB appears physically available, and how do you resolve it?

Differentiating between physical VRAM exhaustion and virtual memory address fragmentation in PyTorch, and tuning the PyTorch CUDA caching allocator for zero-OOM execution.

#CUDA OOM #PyTorch #Memory Allocator #GPU Memory #VRAM #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"The PyTorch CUDA Caching Allocator avoids frequent, expensive calls to `cudaMalloc` and `cudaFree` by maintaining an internal pool of allocated memory blocks. However, alternating allocations of varying tensor sizes cause virtual memory address fragmentation: memory is physically present in the pool, but no single contiguous block is large enough to satisfy the requested tensor allocation, triggering a fatal OOM error."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Analyze Allocated vs Reserved vs Free GPU Memory

Deconstruct the error message: Total VRAM is 79.15 GiB. Allocated (active tensors) is 42.30 GiB. Reserved (cached by PyTorch allocator) is 74.15 GiB (42.30 + 31.85). Free according to CUDA driver is only 4.00 GiB. PyTorch tried to allocate 512 MiB: it searched its 31.85 GiB cache, found no contiguous 512 MiB free block due to fragmentation, called cudaMalloc for a new block, and failed because the driver only had 4.00 GiB left and system allocations prevented it.

# Print memory summary in Python
import torch
print(torch.cuda.memory_summary(device=0, abbreviated=False))
2

Activate PyTorch Expandable Segments to Eliminate Fragmentation

In PyTorch 2.1+, set `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. This enables PyTorch to use CUDA Virtual Memory Management (VMM) APIs (`cuMemCreate`, `cuMemMap`). Instead of requiring physically contiguous memory blocks in the caching allocator, PyTorch dynamically maps virtual memory pages, effectively eradicating allocator fragmentation.

# Set environment variable in Dockerfile or Kubernetes Pod spec
env:
  - name: PYTORCH_CUDA_ALLOC_CONF
    value: 'expandable_segments:True,max_split_size_mb:128'
Advertisement
3

Implement Activation Checkpointing and Garbage Collection Hooks

For deep transformer layers, enable activation checkpointing (gradient checkpointing) using `torch.utils.checkpoint.checkpoint`. This drops intermediate activations during the forward pass and recomputes them during the backward pass, reducing peak VRAM requirements by 60-70% at the cost of ~20-30% additional compute.

from torch.utils.checkpoint import checkpoint

class TransformerBlock(torch.nn.Module):
    def forward(self, x):
        # Recomputes activations on backward pass rather than holding in VRAM
        return checkpoint(self._forward_impl, x, use_reentrant=False)
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"CUDA OOM often results from caching allocator fragmentation rather than true physical VRAM exhaustion. Setting `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` leverages CUDA virtual memory mapping to eliminate fragmentation without manual cache clears."
⚡ 60-Second Elevator Pitch Talking Points
  • An OOM with gigabytes of free memory is almost always caused by PyTorch caching allocator fragmentation, where no contiguous block fits the request.
  • We enable `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which uses CUDA Virtual Memory APIs to map fragmented physical chunks into contiguous virtual addresses.
  • Combined with activation checkpointing, this completely eliminates fragmentation OOMs and allows training 30% larger batch sizes on the same hardware.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →