⚡ 50 Live Scenarios
🎯 STAR Method Answers📋 60s Elevator Pitches
The explosive growth of generative AI, large language models (LLMs), and foundation model pre-training has elevated AI Infrastructure Engineers into some of the most critical and highly compensated roles in modern tech. Building infrastructure for AI is fundamentally different from traditional stateless web services: it demands extreme hardware synchronization, high-bandwidth lossless networking, and surgical management of scarce GPU memory. In senior AI infrastructure loops, hiring managers evaluate your deep understanding of the full hardware and software stack: from NVIDIA GPU Operator and Kubelet Topology Manager NUMA node pinning, to InfiniBand and RoCE v2 lossless networking with Priority Flow Control (PFC) and ECN buffer tuning. Candidates are expected to debug silent distributed training failures—such as multi-node NCCL AllReduce timeouts, CUDA caching allocator virtual memory fragmentation, and GPU Xid error traps—as well as configure gang-scheduling engines like Volcano and Kueue to prevent distributed training deadlocks. On the inference side, interviewers probe your mastery over high-throughput LLM serving runtimes: vLLM PagedAttention KV cache paging, Triton Inference Server dynamic batching, continuous iteration-level batching, dynamic multi-LoRA swapping, and disaggregated prefill-decode architectures over RDMA. Our scenario-based AI/ML infrastructure questions prepare you for Staff and Senior AI Platform Engineer rounds with real-world production configurations, NCCL debugging runbooks, and GPU cluster optimization strategies.
Get 1 AI/ML Infrastructure & GPU interview question in your inbox every week
Join 14,000+ engineers leveling up their cloud and platform interview game. Subscribe to get our weekly deep-dive scenario plus instant access to the Top 50 Kubernetes Interview Questions & Incident Runbooks PDF.
Architectural breakdown and troubleshooting of the NVIDIA Kubernetes Device Plugin, NVML communication, Kubelet Topology Manager, and GPU device injection into ...
Root cause isolation and systematic debugging of NCCL ring initialization timeouts, watchdog hangs, and silent socket deadlocks during multi-node LLM training r...
Differentiating between physical VRAM exhaustion and virtual memory address fragmentation in PyTorch, and tuning the PyTorch CUDA caching allocator for zero-OOM...
Engineering sub-2-minute GPU node provisioning with Karpenter, heterogeneous fallback pools, and automated local NVMe RAID0 initialization for multi-gigabyte mo...
Tuning NVIDIA Triton Inference Server dynamic batching algorithms, queue delay windows, and instance group concurrency to maximize GPU throughput while satisfyi...
Deep dive into vLLM's PagedAttention algorithm, dynamic virtual memory paging of transformer Key-Value caches, continuous iteration-level batching, and `gp...
Deep architectural analysis of InfiniBand vs RoCE v2 fabrics, resolving Priority Flow Control (PFC) deadlocks, storm propagation, and ECN buffer marking in larg...
Deploying and managing KubeRay on Kubernetes, optimizing the Plasma Shared-Memory Object Store, actor scheduling, and triaging worker pod OOMKills without crash...
Engineering high-throughput, non-blocking distributed checkpointing from multi-node GPU clusters to cloud object storage (S3/GCS) without training loop stalls....
Deploying the Kubeflow Training Operator (PyTorchJob), implementing all-or-nothing gang scheduling with Volcano or Kueue, and eradicating distributed job resour...
Architectural evaluation and implementation of multi-terabyte dataset streaming from cloud storage into PyTorch DataLoaders without local disk capacity exhausti...
Comparing Zero Redundancy Optimizer stages (ZeRO-1/2/3) against PyTorch Fully Sharded Data Parallel (FSDP), communication overhead trade-offs, and host CPU/NVMe...
Deep architectural comparison of Slurm vs Kubernetes for AI training workloads, job queuing models, and deploying containerized HPC stacks with Enroot, Pyxis, a...
Architectural selection between Tensor Parallelism (TP) and Pipeline Parallelism (PP) in vLLM/SGLang, evaluating NVLink vs PCIe bandwidth constraints and intra-...
Designing multi-stage machine learning inference pipelines in Triton Inference Server using zero-copy Model Ensembles and dynamic Python Business Logic Scriptin...
Hardening enterprise MLflow tracking and model registry services on Kubernetes with OIDC authentication, scoped AWS IAM IRSA credentials, and secure S3 artifact...
Deep architectural and algorithmic analysis comparing PyTorch Distributed Data Parallel (DDP) against Fully Sharded Data Parallel (FSDP), communication primitiv...
Deploying and configuring NVIDIA Data Center GPU Manager (DCGM) Exporter on Kubernetes, instrumenting critical Prometheus metrics for SM utilization, NVLink ban...
Engineering autoscaling policies for Ray Serve clusters on Kubernetes, driving horizontal replica scaling using request queue depth, pending actors, and inter-t...
Operating the NVIDIA GPU Operator on Kubernetes, leveraging CUDA Forward Compatibility to run modern CUDA 12.x workloads on older enterprise host drivers withou...
Architectural comparison and production benchmark selection between Hugging Face Text Generation Inference (TGI) and vLLM for high-concurrency enterprise LLM se...
Architecting a unified hybrid GPU orchestration plane spanning hyperscalers (AWS/GCP) and specialized AI clouds (CoreWeave/Lambda Labs) with encrypted mesh netw...
Engineering high-density multi-tenant LLM inference platforms that dynamically serve hundreds of fine-tuned LoRA adapters over a single shared base model with z...
Using NVIDIA Triton Model Analyzer to systematically profile model throughput against p99 latency constraints, automate batch size selection, and execute data-d...
#Triton Model Analyzer#Capacity Planning#Profiling
Deep triage of NVIDIA HGX H100 SXM5 systems, resolving Fabric Manager service crashes, NVSwitch link degradation, and PCIe-to-NVSwitch initialization failures i...
Engineering sub-20ms content moderation and jailbreak prevention pipelines using NVIDIA NeMo Guardrails and lightweight classification models integrated with st...
Configuring Kubernetes Topology Manager, Node Feature Discovery (NFD), and CPU Manager to guarantee strict hardware locality between GPUs, CPU sockets, and Mell...
Architectural evaluation framework for selecting between Ray Serve, NVIDIA Triton Inference Server, and vLLM across LLMs, vision models, multi-modal pipelines, ...
Root cause isolation and tuning of PyTorch DataLoader bottlenecks causing low GPU compute utilization, including worker process tuning, host pinned memory, and ...
Diagnosing and fixing CUDA runtime host driver mismatches, missing SM compute architecture compilation targets, and long JIT compilation freezes during PyTorch ...
Defining, measuring, and optimizing Service Level Objectives (SLOs) for generative AI: Time-To-First-Token (TTFT), Inter-Token Latency (ITL), and Time-Per-Outpu...
Evaluating and deploying quantized LLMs in production: comparing native FP8 (Ada/Hopper) against INT4/INT8 (AWQ vs GPTQ), analyzing compute vs memory-bound regi...
Operating Kubeflow Katib for automated large-scale hyperparameter optimization on Kubernetes, implementing Bayesian search algorithms, Median Stopping early ter...
Analyzing the performance impact of network encryption (IPsec / MACsec) on GPUDirect RDMA collective communications, and architecting secure enclaves for propri...
Deep architectural analysis of disaggregated LLM serving: separating compute-bound prompt prefill nodes from memory-bandwidth-bound token decode nodes over high...
Key incident runbooks, interview talking points, and architecture tradeoffs.
How does the NVIDIA Device Plugin discover, advertise, and allocate GPUs to Kubernetes pods via NVML and Kubelet Device Manager, and how do you configure Kubelet Topology Manager to ensure strict NUMA node alignment between GPUs, CPUs, and high-speed NICs?
Kubernetes core does not natively understand PCI topology or GPU hardware. The NVIDIA Kubernetes Device Plugin interfaces with the NVIDIA Management Library (NVML) on the host and registers extended resources (`nvidia.com/gpu`) with Kubelet's Device Manager gRPC service. Without proper NUMA and Topology Manager configuration, GPUs on socket 1 may be bound to CPUs on socket 0, causing severe PCIe interconnect contention and throughput penalties.
Key Architectural Takeaway: The NVIDIA Device Plugin uses NVML to discover and advertise GPU UUIDs to Kubelet via gRPC, and container runtimes inject hardware via CDI. Kubelet Topology Manager with single-numa-node policy is critical to eliminate cross-socket PCIe bandwidth bottlenecks.
A 64-node PyTorch 70B parameter training job crashes after 3 hours with 'NCCL watchdog thread detected timeout: WorkNCCL(OpId=412, Timeout(s)=1800)'. Walk through the step-by-step diagnostic workflow to isolate whether the failure was caused by GPU hardware failure, network packet drop, or PyTorch rank desynchronization.
NCCL (NVIDIA Collective Communications Library) coordinates collective tensor primitives (AllReduce, AllGather, ReduceScatter) across thousands of GPUs. Because all ranks must participate synchronously in collective operations, a single dropped packet, slow rank, or silent GPU freeze causes all other ranks to block indefinitely until the NCCL watchdog timeout expires. Pinpointing the faulty rank requires structured telemetry analysis.
Key Architectural Takeaway: NCCL timeouts mask the true culprit because healthy nodes block waiting on the single hung rank. Always enable PyTorch NCCL Flight Recorder, correlate timestamps against host kernel Xid errors, and check RDMA PFC/pause counters for fabric packet loss.
A training pod crashes with 'torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB (GPU 0; 79.15 GiB total capacity; 42.30 GiB already allocated; 31.85 GiB reserved in PyTorch; 4.00 GiB free)'. Why did this OOM occur when 36.85 GiB appears physically available, and how do you resolve it?
The PyTorch CUDA Caching Allocator avoids frequent, expensive calls to `cudaMalloc` and `cudaFree` by maintaining an internal pool of allocated memory blocks. However, alternating allocations of varying tensor sizes cause virtual memory address fragmentation: memory is physically present in the pool, but no single contiguous block is large enough to satisfy the requested tensor allocation, triggering a fatal OOM error.
Key Architectural Takeaway: CUDA OOM often results from caching allocator fragmentation rather than true physical VRAM exhaustion. Setting `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` leverages CUDA virtual memory mapping to eliminate fragmentation without manual cache clears.
Standard Cluster Autoscaler takes 12-15 minutes to provision a p4de.24xlarge node and pull a 140GB model image. How do you design Karpenter NodePools with heterogeneous instance types, pre-warmed AMIs, and automated NVMe instance store RAID0 mounting to slash scale-up time to under 2 minutes?
GPU instances (such as AWS `g5`, `p4d`, `p4de`, `p5`) are scarce and take significant time to launch. Furthermore, pulling massive container images or downloading 100GB+ LLM checkpoints over EBS during pod startup ruins autoscaling responsiveness. Karpenter provisions compute directly to pending pod requirements, but requires custom userdata and storage architectures to deliver rapid startup times.
Key Architectural Takeaway: Achieving sub-2-minute GPU scale-up requires Karpenter's direct EC2 API provisioning across heterogeneous instance pools, combined with automated local NVMe RAID0 creation in userdata to serve as a high-speed shared model weights cache.
How does Triton Inference Server's dynamic batcher aggregate asynchronous individual inference requests, and how do you calculate optimal values for `max_batch_size`, `max_queue_delay_microseconds`, and `instance_group` counts under a 50ms p99 SLA?
Inference requests arrive individually from microservices, but GPUs achieve peak arithmetic intensity and hardware utilization only when processing batched tensors. NVIDIA Triton Inference Server provides a server-side dynamic batcher that intercepts individual requests and combines them into batches without requiring client-side batching logic.
Key Architectural Takeaway: Triton's dynamic batcher aggregates concurrent requests within a `max_queue_delay_microseconds` window. Tuning batch sizes alongside `instance_group` concurrency via Model Analyzer maximizes GPU saturation while respecting tight latency budgets.
🌐
Explore Related DevOps & Cloud Domains
Cross-train across interconnected systems for senior and staff infrastructure rounds.
Master AI/ML Infrastructure & GPU & Platform Engineering In Production
Join 14,000+ engineers leveling up their cloud and platform interview game. Subscribe to get our weekly deep-dive scenario plus instant access to the Top 50 Kubernetes Interview Questions & Incident Runbooks PDF.