⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 40 of 50 in AI/ML Infrastructure & GPU
Senior AI Infrastructure Engineer AI/ML Infrastructure Distributed Training Distributed Training
🎯 Target Role / Context: Senior AI Infrastructure Engineer optimizing deep learning training throughput and cluster efficiency.

Q: An 8x A100 training job shows only 25% GPU SM utilization while CPUs are pegged at 100%. Why is the GPU starved for data, how do you tune `num_workers`, `pin_memory=True`, and `persistent_workers`, and why does inadequate `/dev/shm` crash PyTorch DataLoaders?

Root cause isolation and tuning of PyTorch DataLoader bottlenecks causing low GPU compute utilization, including worker process tuning, host pinned memory, and `/dev/shm` IPC sizing.

#DataLoader #GPU Starvation #pin_memory #Shared Memory #PyTorch #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"When training deep neural networks, GPUs can process tensors orders of magnitude faster than a CPU can load, decompress, and augment raw images or tokens from disk. If the data pipeline cannot feed tensors to the GPU fast enough, the GPU sits idle during every step waiting on the next batch, wasting expensive compute."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Profile the Training Step Time: Data Loading vs Forward/Backward Pass

Instrument the PyTorch training loop using PyTorch Profiler (`torch.profiler`). Analyze the step breakdown: if `DataLoader:next` accounts for 70% of step time while `cudaStreamSynchronize` is small, the pipeline is severely CPU-starved and data-bound.

with torch.profiler.profile(
    activities=[torch.profiler.ProfilerActivity.CPU, torch.profiler.ProfilerActivity.CUDA],
    schedule=torch.profiler.schedule(wait=1, warmup=1, active=3, repeat=2),
    on_trace_ready=torch.profiler.tensorboard_trace_handler('/tmp/log/profiler')
) as prof:
    for batch in dataloader:
        train_step(batch)
        prof.step()
2

Tune PyTorch DataLoader Concurrency Parameters

Configure DataLoader settings: (1) Set `num_workers = 4` to `8` per GPU (avoid setting higher than available physical CPU cores). (2) Enable `pin_memory=True`: this allocates tensors in host page-locked (pinned) memory, allowing fast asynchronous DMA copy directly into GPU VRAM. (3) Set `persistent_workers=True`: this keeps DataLoader worker processes alive between epochs, avoiding expensive process respawn and dataset reinitialization overhead.

dataloader = DataLoader(
    dataset,
    batch_size=128,
    num_workers=8,
    pin_memory=True,
    persistent_workers=True,
    prefetch_factor=4
)
Advertisement
3

Size Kubernetes /dev/shm to Prevent IPC Bus Errors

PyTorch DataLoader workers pass tensors to the main training process via shared memory IPC (`/dev/shm`). In Kubernetes, the default `/dev/shm` size is only 64MB! Once batches exceed 64MB, workers crash with `RuntimeError: DataLoader worker (pid XXX) is killed by signal: Bus error`. Always mount an `emptyDir` with `medium: Memory` sized to 32GB+ for `/dev/shm` in the pod spec.

volumes:
- name: dshm
  emptyDir:
    medium: Memory
    sizeLimit: 64Gi
volumeMounts:
- mountPath: /dev/shm
  name: dshm
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Low GPU utilization during training is frequently caused by CPU preprocessing bottlenecks. Enable `pin_memory=True`, `persistent_workers=True`, tune `num_workers`, and ensure `/dev/shm` is mounted with sufficient memory to prevent IPC bus crashes."
⚡ 60-Second Elevator Pitch Talking Points
  • GPUs sitting at 25% utilization is usually caused by slow CPU data preprocessing, not slow model code.
  • We tune PyTorch DataLoaders with `pin_memory=True` for DMA transfers, `persistent_workers=True` to eliminate epoch restarts, and prefetch factor 4.
  • We size Kubernetes `/dev/shm` to 64GB in RAM, preventing the classic 'Bus error' crash when multiple workers transfer large tensor batches.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →