3.3 Monitoring GPUs

NCA-AIIO · AI Operations (22% of the exam) · Official objective: “Articulate the key measures and criteria related to monitoring GPUs.”

The DCGM metrics that matter and how to export them to Prometheus.

Key points

  1. Field 311 counts double-bit volatile ECC errors. A double-bit error cannot be corrected by ECC, so any non-zero value is significant. NVIDIA's Xid Catalog logs uncorrectable errors as Xid 48, with RESET_GPU and RUN_FIELDDIAG as the actions when it appears alone.

    What NVIDIA says (3)

    “DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”

    — DCGM Field IDs

    “This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

    — NVIDIA Xid Catalog

    “WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”

    — NVIDIA Xid Catalog

  2. When a GPU lowers its clocks to protect itself, the reason is recorded as a clock event. DCGM field 112 reports the current clock event reasons as a bitmask, so you can see whether heat, power or something else caused the drop. Field 100 shows the clock value, but not the reason.

    What NVIDIA says (2)

    “DCGM_FI_DEV_CLOCKS_EVENT_REASONS 112 Current clock event reasons (bitmask of DCGM_CLOCKS_EVENT_REASON_*)”

    — DCGM Field IDs

    “DCGM_FI_DEV_SM_CLOCK 100 SM clock for the device.”

    — DCGM Field IDs

  3. DCGM (Data Center GPU Manager) exposes GPU telemetry as numbered fields. Field 310, DCGM_FI_DEV_ECC_SBE_VOL_TOTAL, is the total of single-bit volatile ECC errors. Volatile means the count since the driver last loaded. Single-bit errors are corrected, so the trend matters more than any one value.

    What NVIDIA says (1)

    “DCGM_FI_DEV_ECC_SBE_VOL_TOTAL 310 Total single bit volatile ECC errors.”

    — DCGM Field IDs

  4. Prometheus is an open-source monitoring system that scrapes metrics over HTTP. DCGM Exporter exposes GPU metrics at an HTTP /metrics endpoint for tools such as Prometheus. Its default listen address is :9400.

    What NVIDIA says (2)

    “DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”

    — NVIDIA DCGM Exporter

    “Address of listening http server. Default: “:9400””

    — NVIDIA DCGM Exporter

  5. DCGM (Data Center GPU Manager) is NVIDIA's suite for managing and monitoring data center GPUs. NVIDIA says it includes active health monitoring, diagnostics, system alerts and governance policies such as power and clock management. Teams can use it alone or integrate it into cluster management and monitoring tools.

    What NVIDIA says (5)

    “Active Health Checks (GPU subsystems)”

    — DCGM Feature Overview

    “GPU Diagnostics (Diagnostic Levels - 1, 2, 3)”

    — DCGM Feature Overview

    “It includes active health monitoring, comprehensive diagnostics, system alerts, and governance policies including power and clock management.”

    — NVIDIA DCGM (developer page)

    “The core DCGM library can be run as a standalone process or be loaded by an agent as a shared library.”

    — DCGM User Guide: Getting Started

    “Infrastructure teams can use it standalone and in addition easily integrate it into cluster management tools, resource scheduling, and monitoring products from NVIDIA partners.”

    — NVIDIA DCGM (developer page)

  6. A streaming multiprocessor (SM) is one of the many processing blocks inside a GPU. DCGM defines graphics engine activity as the fraction of time any portion of the graphics or compute engines were active. So it can be high while most SMs are idle. DCGM's SM activity metric is the fraction of time at least one warp was active, averaged over all SMs. A warp is a group of threads the SM runs together. That metric shows how much of the chip is really working.

    What NVIDIA says (2)

    “Graphics Engine Activity The fraction of time any portion of the graphics or compute engines were active.”

    — DCGM Feature Overview

    “SM Activity The fraction of time at least one warp was active on a multiprocessor, averaged over all multiprocessors.”

    — DCGM Feature Overview

Key terms

Try it

Sample question

DCGM field 311 (DCGM_FI_DEV_ECC_DBE_VOL_TOTAL) reads 1 on GPU 3. What has happened?

Show the answer

Answer: The GPU recorded a double-bit ECC error, which is uncorrectable. The Xid Catalog treats it as Xid 48: reset the GPU and run field diagnostics.

Field 311 counts double-bit volatile ECC errors. A double-bit error cannot be corrected by ECC, so any non-zero value is significant. NVIDIA's Xid Catalog logs uncorrectable errors as Xid 48, with RESET_GPU and RUN_FIELDDIAG as the actions when it appears alone.

What NVIDIA says (3)

“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”

— DCGM Field IDs

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— NVIDIA Xid Catalog

“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”

— NVIDIA Xid Catalog

Practice 3.3 (6 questions) Full AI Operations guide

← 3.2 Orchestration and job scheduling · 3.4 Virtualizing GPUs →