Q: A container built with CUDA 12.2 and PyTorch 2.3 freezes for 10 minutes on startup or crashes with 'RuntimeError: CUDA error: no kernel image is available for execution on the device'. How do you troubleshoot CUDA compute architecture targets, PTX JIT compilation, and driver compatibility?
Diagnosing and fixing CUDA runtime host driver mismatches, missing SM compute architecture compilation targets, and long JIT compilation freezes during PyTorch startup.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Decode 'No kernel image is available for execution on the device'
This error occurs when the PyTorch binary or custom CUDA extension was not compiled for the target GPU's Compute Capability. For example, running a wheel compiled for `sm_75,sm_80` on an NVIDIA H100 (`sm_90`) will fail because no matching cubin exists and no forward-compatible PTX was embedded in the fatbinary.
# Check GPU Compute Capability
nvidia-smi --query-gpu=name,compute_cap --format=csv
# Check PyTorch supported architectures in Python
python3 -c 'import torch; print(torch.cuda.get_arch_list())'
# Output: ['sm_70', 'sm_75', 'sm_80', 'sm_86', 'sm_90']
Diagnose PTX JIT Startup Freezes and Configure TORCH_CUDA_ARCH_LIST
If PyTorch lacks the exact cubin but has PTX, the driver runs JIT compilation on the first kernel invocation, causing a multi-minute freeze that often triggers Kubernetes liveness probe timeouts. When building custom CUDA kernels (e.g., FlashAttention, vLLM, DeepSpeed), always set `TORCH_CUDA_ARCH_LIST` explicitly to compile native cubin binaries for your exact target architectures.
# Build with native cubins for A100 (8.0) and H100 (9.0)
export TORCH_CUDA_ARCH_LIST="8.0 9.0+PTX"
python3 setup.py install
Enable CUDA JIT Caching to Amortize Compilation Overhead
If runtime JIT compilation is unavoidable, configure the CUDA JIT cache. By default, the cache is limited to a small size. Set `CUDA_CACHE_MAXSIZE=4294967296` (4GB) and mount a persistent cache directory (`CUDA_CACHE_PATH=/mnt/cache/cuda`) so subsequent pod restarts skip JIT recompilation entirely.
env:
- name: CUDA_CACHE_PATH
value: /var/cache/cuda
- name: CUDA_CACHE_MAXSIZE
value: '4294967296'
- The error 'no kernel image available' means your PyTorch binary lacks pre-compiled code for your GPU's Compute Capability architecture.
- Falling back to PTX JIT compilation freezes pod startup for 10 minutes, failing Kubernetes health probes.
- We compile custom kernels with `TORCH_CUDA_ARCH_LIST="8.0 9.0+PTX"` and configure a persistent 4GB CUDA JIT cache, ensuring instant pod startup across both A100 and H100 nodes.