Q: An 8x A100 training job shows only 25% GPU SM utilization while CPUs are pegged at 100%. Why is the GPU starved for data, how do you tune `num_workers`, `pin_memory=True`, and `persistent_workers`, and why does inadequate `/dev/shm` crash PyTorch DataLoaders?
Root cause isolation and tuning of PyTorch DataLoader bottlenecks causing low GPU compute utilization, including worker process tuning, host pinned memory, and `/dev/shm` IPC sizing.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Profile the Training Step Time: Data Loading vs Forward/Backward Pass
Instrument the PyTorch training loop using PyTorch Profiler (`torch.profiler`). Analyze the step breakdown: if `DataLoader:next` accounts for 70% of step time while `cudaStreamSynchronize` is small, the pipeline is severely CPU-starved and data-bound.
with torch.profiler.profile(
activities=[torch.profiler.ProfilerActivity.CPU, torch.profiler.ProfilerActivity.CUDA],
schedule=torch.profiler.schedule(wait=1, warmup=1, active=3, repeat=2),
on_trace_ready=torch.profiler.tensorboard_trace_handler('/tmp/log/profiler')
) as prof:
for batch in dataloader:
train_step(batch)
prof.step()
Tune PyTorch DataLoader Concurrency Parameters
Configure DataLoader settings: (1) Set `num_workers = 4` to `8` per GPU (avoid setting higher than available physical CPU cores). (2) Enable `pin_memory=True`: this allocates tensors in host page-locked (pinned) memory, allowing fast asynchronous DMA copy directly into GPU VRAM. (3) Set `persistent_workers=True`: this keeps DataLoader worker processes alive between epochs, avoiding expensive process respawn and dataset reinitialization overhead.
dataloader = DataLoader(
dataset,
batch_size=128,
num_workers=8,
pin_memory=True,
persistent_workers=True,
prefetch_factor=4
)
Size Kubernetes /dev/shm to Prevent IPC Bus Errors
PyTorch DataLoader workers pass tensors to the main training process via shared memory IPC (`/dev/shm`). In Kubernetes, the default `/dev/shm` size is only 64MB! Once batches exceed 64MB, workers crash with `RuntimeError: DataLoader worker (pid XXX) is killed by signal: Bus error`. Always mount an `emptyDir` with `medium: Memory` sized to 32GB+ for `/dev/shm` in the pod spec.
volumes:
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 64Gi
volumeMounts:
- mountPath: /dev/shm
name: dshm
- GPUs sitting at 25% utilization is usually caused by slow CPU data preprocessing, not slow model code.
- We tune PyTorch DataLoaders with `pin_memory=True` for DMA transfers, `persistent_workers=True` to eliminate epoch restarts, and prefetch factor 4.
- We size Kubernetes `/dev/shm` to 64GB in RAM, preventing the classic 'Bus error' crash when multiple workers transfer large tensor batches.