3.3 Monitoring GPUs
The DCGM metrics that matter and how to export them to Prometheus.
Key points
Field 311 counts double-bit volatile ECC errors. A double-bit error cannot be corrected by ECC, so any non-zero value is significant. NVIDIA's Xid Catalog logs uncorrectable errors as Xid 48, with RESET_GPU and RUN_FIELDDIAG as the actions when it appears alone.
What NVIDIA says (3)
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”
When a GPU lowers its clocks to protect itself, the reason is recorded as a clock event. DCGM field 112 reports the current clock event reasons as a bitmask, so you can see whether heat, power or something else caused the drop. Field 100 shows the clock value, but not the reason.
What NVIDIA says (2)
“DCGM_FI_DEV_CLOCKS_EVENT_REASONS 112 Current clock event reasons (bitmask of DCGM_CLOCKS_EVENT_REASON_*)”
“DCGM_FI_DEV_SM_CLOCK 100 SM clock for the device.”
DCGM (Data Center GPU Manager) exposes GPU telemetry as numbered fields. Field 310, DCGM_FI_DEV_ECC_SBE_VOL_TOTAL, is the total of single-bit volatile ECC errors. Volatile means the count since the driver last loaded. Single-bit errors are corrected, so the trend matters more than any one value.
What NVIDIA says (1)
“DCGM_FI_DEV_ECC_SBE_VOL_TOTAL 310 Total single bit volatile ECC errors.”
Prometheus is an open-source monitoring system that scrapes metrics over HTTP. DCGM Exporter exposes GPU metrics at an HTTP /metrics endpoint for tools such as Prometheus. Its default listen address is :9400.
What NVIDIA says (2)
“DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”
“Address of listening http server. Default: “:9400””
DCGM (Data Center GPU Manager) is NVIDIA's suite for managing and monitoring data center GPUs. NVIDIA says it includes active health monitoring, diagnostics, system alerts and governance policies such as power and clock management. Teams can use it alone or integrate it into cluster management and monitoring tools.
What NVIDIA says (5)
“Active Health Checks (GPU subsystems)”
“GPU Diagnostics (Diagnostic Levels - 1, 2, 3)”
“It includes active health monitoring, comprehensive diagnostics, system alerts, and governance policies including power and clock management.”
“The core DCGM library can be run as a standalone process or be loaded by an agent as a shared library.”
“Infrastructure teams can use it standalone and in addition easily integrate it into cluster management tools, resource scheduling, and monitoring products from NVIDIA partners.”
A streaming multiprocessor (SM) is one of the many processing blocks inside a GPU. DCGM defines graphics engine activity as the fraction of time any portion of the graphics or compute engines were active. So it can be high while most SMs are idle. DCGM's SM activity metric is the fraction of time at least one warp was active, averaged over all SMs. A warp is a group of threads the SM runs together. That metric shows how much of the chip is really working.
What NVIDIA says (2)
“Graphics Engine Activity The fraction of time any portion of the graphics or compute engines were active.”
“SM Activity The fraction of time at least one warp was active on a multiprocessor, averaged over all multiprocessors.”
Key terms
- ECC: Memory protection that fixes single-bit errors and detects double-bit errors.
- DCGM: NVIDIA tooling for GPU health monitoring, diagnostics, alerts and policies across a cluster.
- DCGM Exporter: A service that publishes DCGM GPU metrics on an HTTP /metrics endpoint for Prometheus.
- Prometheus: An open-source monitoring system that collects metrics, such as DCGM Exporter GPU metrics, from /metrics endpoints.
Try it
Sample question
DCGM field 311 (DCGM_FI_DEV_ECC_DBE_VOL_TOTAL) reads 1 on GPU 3. What has happened?
Show the answer
Answer: The GPU recorded a double-bit ECC error, which is uncorrectable. The Xid Catalog treats it as Xid 48: reset the GPU and run field diagnostics.
Field 311 counts double-bit volatile ECC errors. A double-bit error cannot be corrected by ECC, so any non-zero value is significant. NVIDIA's Xid Catalog logs uncorrectable errors as Xid 48, with RESET_GPU and RUN_FIELDDIAG as the actions when it appears alone.
What NVIDIA says (3)
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”
Practice 3.3 (6 questions) Full AI Operations guide
← 3.2 Orchestration and job scheduling · 3.4 Virtualizing GPUs →