⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All AI/ML Infrastructure & GPU Interview Questions Scenario 21 of 50 in AI/ML Infrastructure & GPU
Senior AI Infrastructure Engineer AI/ML Infrastructure Hardware Reliability & Observability GPU Observability
🎯 Target Role / Context: Senior AI Infrastructure Engineer creating production observability dashboards and alert rules for deep learning clusters.

Q: Which specific Prometheus metrics from NVIDIA DCGM Exporter distinguish true GPU compute utilization from memory-copy stalls, and how do you detect NVLink bandwidth degradation and thermal clock throttling across thousands of GPUs?

Deploying and configuring NVIDIA Data Center GPU Manager (DCGM) Exporter on Kubernetes, instrumenting critical Prometheus metrics for SM utilization, NVLink bandwidth, and thermal/power throttling.

#DCGM #NVIDIA DCGM Exporter #Prometheus #GPU Telemetry #Grafana #AI/ML Infra
🎙️ Candidate Opening & Architectural Context
"Standard Kubernetes metrics (CPU and RAM) tell you nothing about GPU workloads. Furthermore, the generic `DCGM_FI_DEV_GPU_UTIL` metric can be deeply misleading: it reports 100% if a GPU is merely waiting on memory transfers from host RAM while compute cores sit completely idle. Engineering robust GPU observability requires capturing low-level hardware counters exposed by NVIDIA Data Center GPU Manager (DCGM)."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Deploy NVIDIA DCGM Exporter DaemonSet with Custom Profiling Counters

Deploy DCGM Exporter via the NVIDIA GPU Operator or Helm chart. By default, DCGM Exporter uses a basic metrics list. Upgrade to the profiling metrics configuration (`dcp-metrics-included.csv`), enabling hardware performance counters for Streaming Multiprocessor (SM) active cycles, tensor core activity, and memory bus utilization.

# helm values for dcgm-exporter
serviceMonitor:
  enabled: true
  interval: 15s
arguments:
  - "-f"
  - "/etc/dcgm-exporter/dcp-metrics-included.csv"
2

Differentiate Compute vs Memory Stalls: SM Utilization vs Tensor Pipe Active

To measure true model compute efficiency, correlate: `DCGM_FI_PROF_SM_ACTIVE` (fraction of time SMs have at least one warp executing) and `DCGM_FI_PROF_PIPE_TENSOR_ACTIVE` (fraction of time Tensor Cores are active, indicating FP16/BF16/FP8 matrix multiplications) against `DCGM_FI_PROF_DRAM_ACTIVE` (memory bandwidth saturation). If `SM_ACTIVE` is low while `DRAM_ACTIVE` is 90%+, the workload is memory-bandwidth bound rather than compute bound.

# Prometheus PromQL for true Tensor Core utilization percentage
avg(DCGM_FI_PROF_PIPE_TENSOR_ACTIVE) by (pod, namespace) * 100
Advertisement
3

Detect Hardware Throttling and NVLink Bottlenecks

Instrument alerts for GPU hardware throttling: monitor `DCGM_FI_DEV_CLOCK_THROTTLE_REASONS` bitmasks. Bit 0x0000000000000008 indicates thermal throttling (temperatures exceeding thermal limits), and bit 0x0000000000000004 indicates power brake throttling. Track `DCGM_FI_PROF_NVLINK_TX_BYTES` and `DCGM_FI_PROF_NVLINK_RX_BYTES` to detect imbalanced tensor parallel communication across GPU ranks.

# Alert rule: GPU Thermal or Power Throttling detected
- alert: GPUClockThrottled
  expr: DCGM_FI_DEV_CLOCK_THROTTLE_REASONS > 0
  for: 2m
  labels:
    severity: warning
  annotations:
    summary: "GPU {{ $labels.gpu }} on {{ $labels.node }} is throttling clock frequencies"
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Never rely solely on generic GPU utilization. Monitor `DCGM_FI_PROF_PIPE_TENSOR_ACTIVE` to measure real Tensor Core math execution, track `DRAM_ACTIVE` for memory bus saturation, and alert on `CLOCK_THROTTLE_REASONS` to catch thermal and power degradation."
⚡ 60-Second Elevator Pitch Talking Points
  • Basic `nvidia-smi` utilization is misleading because waiting on memory transfers shows up as 100% utilized.
  • We configure DCGM Exporter with profiling metrics to track `SM_ACTIVE` and `PIPE_TENSOR_ACTIVE`, revealing exactly how much time GPUs spend crunching matrix math.
  • We alert on `CLOCK_THROTTLE_REASONS` and monitor NVLink TX/RX byte throughput to catch thermal throttling and interconnect bottlenecks before they ruin training schedules.
Advertisement
Want more AI/ML Infrastructure & GPU scenarios?
Explore our complete collection of scenario-based AI/ML Infrastructure & GPU interview runbooks.
Browse All AI/ML Infrastructure & GPU Questions →