Q: Which specific Prometheus metrics from NVIDIA DCGM Exporter distinguish true GPU compute utilization from memory-copy stalls, and how do you detect NVLink bandwidth degradation and thermal clock throttling across thousands of GPUs?
Deploying and configuring NVIDIA Data Center GPU Manager (DCGM) Exporter on Kubernetes, instrumenting critical Prometheus metrics for SM utilization, NVLink bandwidth, and thermal/power throttling.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy NVIDIA DCGM Exporter DaemonSet with Custom Profiling Counters
Deploy DCGM Exporter via the NVIDIA GPU Operator or Helm chart. By default, DCGM Exporter uses a basic metrics list. Upgrade to the profiling metrics configuration (`dcp-metrics-included.csv`), enabling hardware performance counters for Streaming Multiprocessor (SM) active cycles, tensor core activity, and memory bus utilization.
# helm values for dcgm-exporter
serviceMonitor:
enabled: true
interval: 15s
arguments:
- "-f"
- "/etc/dcgm-exporter/dcp-metrics-included.csv"
Differentiate Compute vs Memory Stalls: SM Utilization vs Tensor Pipe Active
To measure true model compute efficiency, correlate: `DCGM_FI_PROF_SM_ACTIVE` (fraction of time SMs have at least one warp executing) and `DCGM_FI_PROF_PIPE_TENSOR_ACTIVE` (fraction of time Tensor Cores are active, indicating FP16/BF16/FP8 matrix multiplications) against `DCGM_FI_PROF_DRAM_ACTIVE` (memory bandwidth saturation). If `SM_ACTIVE` is low while `DRAM_ACTIVE` is 90%+, the workload is memory-bandwidth bound rather than compute bound.
# Prometheus PromQL for true Tensor Core utilization percentage
avg(DCGM_FI_PROF_PIPE_TENSOR_ACTIVE) by (pod, namespace) * 100
Detect Hardware Throttling and NVLink Bottlenecks
Instrument alerts for GPU hardware throttling: monitor `DCGM_FI_DEV_CLOCK_THROTTLE_REASONS` bitmasks. Bit 0x0000000000000008 indicates thermal throttling (temperatures exceeding thermal limits), and bit 0x0000000000000004 indicates power brake throttling. Track `DCGM_FI_PROF_NVLINK_TX_BYTES` and `DCGM_FI_PROF_NVLINK_RX_BYTES` to detect imbalanced tensor parallel communication across GPU ranks.
# Alert rule: GPU Thermal or Power Throttling detected
- alert: GPUClockThrottled
expr: DCGM_FI_DEV_CLOCK_THROTTLE_REASONS > 0
for: 2m
labels:
severity: warning
annotations:
summary: "GPU {{ $labels.gpu }} on {{ $labels.node }} is throttling clock frequencies"
- Basic `nvidia-smi` utilization is misleading because waiting on memory transfers shows up as 100% utilized.
- We configure DCGM Exporter with profiling metrics to track `SM_ACTIVE` and `PIPE_TENSOR_ACTIVE`, revealing exactly how much time GPUs spend crunching matrix math.
- We alert on `CLOCK_THROTTLE_REASONS` and monitor NVLink TX/RX byte throughput to catch thermal throttling and interconnect bottlenecks before they ruin training schedules.