⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
🧠

AI/ML Infrastructure, GPU Orchestration & LLM Serving Interview Questions (2026 Edition)

⚡ 50 Live Scenarios 🎯 STAR Method Answers 📋 60s Elevator Pitches

The explosive growth of generative AI, large language models (LLMs), and foundation model pre-training has elevated AI Infrastructure Engineers into some of the most critical and highly compensated roles in modern tech. Building infrastructure for AI is fundamentally different from traditional stateless web services: it demands extreme hardware synchronization, high-bandwidth lossless networking, and surgical management of scarce GPU memory. In senior AI infrastructure loops, hiring managers evaluate your deep understanding of the full hardware and software stack: from NVIDIA GPU Operator and Kubelet Topology Manager NUMA node pinning, to InfiniBand and RoCE v2 lossless networking with Priority Flow Control (PFC) and ECN buffer tuning. Candidates are expected to debug silent distributed training failures—such as multi-node NCCL AllReduce timeouts, CUDA caching allocator virtual memory fragmentation, and GPU Xid error traps—as well as configure gang-scheduling engines like Volcano and Kueue to prevent distributed training deadlocks. On the inference side, interviewers probe your mastery over high-throughput LLM serving runtimes: vLLM PagedAttention KV cache paging, Triton Inference Server dynamic batching, continuous iteration-level batching, dynamic multi-LoRA swapping, and disaggregated prefill-decode architectures over RDMA. Our scenario-based AI/ML infrastructure questions prepare you for Staff and Senior AI Platform Engineer rounds with real-world production configurations, NCCL debugging runbooks, and GPU cluster optimization strategies.

Filter by Subcategory:
Filter by Level:
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

Advertisement

All AI/ML Infrastructure & GPU Scenario Questions (50)

⚡ Practice in Interactive Simulator
Advertisement
Showing 25 of 50 Scenarios
❓

Frequently Asked AI/ML Infrastructure & GPU Interview Questions

Key incident runbooks, interview talking points, and architecture tradeoffs.

How does the NVIDIA Device Plugin discover, advertise, and allocate GPUs to Kubernetes pods via NVML and Kubelet Device Manager, and how do you configure Kubelet Topology Manager to ensure strict NUMA node alignment between GPUs, CPUs, and high-speed NICs?

Kubernetes core does not natively understand PCI topology or GPU hardware. The NVIDIA Kubernetes Device Plugin interfaces with the NVIDIA Management Library (NVML) on the host and registers extended resources (`nvidia.com/gpu`) with Kubelet's Device Manager gRPC service. Without proper NUMA and Topology Manager configuration, GPUs on socket 1 may be bound to CPUs on socket 0, causing severe PCIe interconnect contention and throughput penalties.

Key Architectural Takeaway: The NVIDIA Device Plugin uses NVML to discover and advertise GPU UUIDs to Kubelet via gRPC, and container runtimes inject hardware via CDI. Kubelet Topology Manager with single-numa-node policy is critical to eliminate cross-socket PCIe bandwidth bottlenecks.
A 64-node PyTorch 70B parameter training job crashes after 3 hours with 'NCCL watchdog thread detected timeout: WorkNCCL(OpId=412, Timeout(s)=1800)'. Walk through the step-by-step diagnostic workflow to isolate whether the failure was caused by GPU hardware failure, network packet drop, or PyTorch rank desynchronization.

NCCL (NVIDIA Collective Communications Library) coordinates collective tensor primitives (AllReduce, AllGather, ReduceScatter) across thousands of GPUs. Because all ranks must participate synchronously in collective operations, a single dropped packet, slow rank, or silent GPU freeze causes all other ranks to block indefinitely until the NCCL watchdog timeout expires. Pinpointing the faulty rank requires structured telemetry analysis.

Key Architectural Takeaway: NCCL timeouts mask the true culprit because healthy nodes block waiting on the single hung rank. Always enable PyTorch NCCL Flight Recorder, correlate timestamps against host kernel Xid errors, and check RDMA PFC/pause counters for fabric packet loss.
A training pod crashes with 'torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB (GPU 0; 79.15 GiB total capacity; 42.30 GiB already allocated; 31.85 GiB reserved in PyTorch; 4.00 GiB free)'. Why did this OOM occur when 36.85 GiB appears physically available, and how do you resolve it?

The PyTorch CUDA Caching Allocator avoids frequent, expensive calls to `cudaMalloc` and `cudaFree` by maintaining an internal pool of allocated memory blocks. However, alternating allocations of varying tensor sizes cause virtual memory address fragmentation: memory is physically present in the pool, but no single contiguous block is large enough to satisfy the requested tensor allocation, triggering a fatal OOM error.

Key Architectural Takeaway: CUDA OOM often results from caching allocator fragmentation rather than true physical VRAM exhaustion. Setting `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` leverages CUDA virtual memory mapping to eliminate fragmentation without manual cache clears.
Standard Cluster Autoscaler takes 12-15 minutes to provision a p4de.24xlarge node and pull a 140GB model image. How do you design Karpenter NodePools with heterogeneous instance types, pre-warmed AMIs, and automated NVMe instance store RAID0 mounting to slash scale-up time to under 2 minutes?

GPU instances (such as AWS `g5`, `p4d`, `p4de`, `p5`) are scarce and take significant time to launch. Furthermore, pulling massive container images or downloading 100GB+ LLM checkpoints over EBS during pod startup ruins autoscaling responsiveness. Karpenter provisions compute directly to pending pod requirements, but requires custom userdata and storage architectures to deliver rapid startup times.

Key Architectural Takeaway: Achieving sub-2-minute GPU scale-up requires Karpenter's direct EC2 API provisioning across heterogeneous instance pools, combined with automated local NVMe RAID0 creation in userdata to serve as a high-speed shared model weights cache.
How does Triton Inference Server's dynamic batcher aggregate asynchronous individual inference requests, and how do you calculate optimal values for `max_batch_size`, `max_queue_delay_microseconds`, and `instance_group` counts under a 50ms p99 SLA?

Inference requests arrive individually from microservices, but GPUs achieve peak arithmetic intensity and hardware utilization only when processing batched tensors. NVIDIA Triton Inference Server provides a server-side dynamic batcher that intercepts individual requests and combines them into batches without requiring client-side batching logic.

Key Architectural Takeaway: Triton's dynamic batcher aggregates concurrent requests within a `max_queue_delay_microseconds` window. Tuning batch sizes alongside `instance_group` concurrency via Model Analyzer maximizes GPU saturation while respecting tight latency budgets.
Advertisement