3.1 Data center management and monitoring

NCA-AIIO · AI Operations (22% of the exam) · Official objective: “Describe AI data center management and monitoring essentials.”

Reading Xid errors and choosing the recovery action NVIDIA documents.

Key points

  1. An Xid is an error report that the NVIDIA driver writes to the kernel log. Xid 48 is a double-bit ECC error. ECC (error-correcting code) memory can fix single-bit errors, but not double-bit ones, so this error is uncorrectable. NVIDIA's Xid Catalog lists the recovery for Xid 48 seen alone as RESET_GPU and the follow-up as RUN_FIELDDIAG (run field diagnostics). If Xid 63 or 64 appears with it, the catalog says to drain the node and reset.

    What NVIDIA says (3)

    “48 ROBUST_CHANNEL_GPU_ECC_DBE Double Bit ECC Error”

    — NVIDIA Xid Catalog

    “This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

    — NVIDIA Xid Catalog

    “WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU w/ 63 or 64: DRAIN_AND_RESET Investagatory Action Solo: RUN_FIELDDIAG”

    — NVIDIA Xid Catalog

  2. Xid 79 means the GPU has fallen off the bus. The bus here is PCI Express (PCIe), the link between the GPU and the server's CPU. NVIDIA says the event is logged when the driver tries to reach the GPU over PCIe and finds it is not accessible. It is often caused by hardware failures on the PCIe link, so check system and kernel PCI event logs.

    What NVIDIA says (3)

    “79 ROBUST_CHANNEL_GPU_HAS_FALLEN_OFF_THE_BUS”

    — NVIDIA Xid Catalog

    “This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”

    — NVIDIA Xid Catalog

    “This event is often caused by hardware failures on the PCI Express link causing the GPU to be inaccessible due to the link being brought down.”

    — NVIDIA Xid Catalog

  3. NVLink is NVIDIA's direct GPU-to-GPU link. An NVSwitch is a switch chip that connects many GPUs over NVLink. Xid 74 is logged when the GPU detects a problem with an NVLink connection to another GPU or an NVSwitch. The cause can be at the far end. If one GPU fails, a GPU linked to it may report Xid 74 simply because the link went down. For several error bits, NVIDIA's workflow says to check link mechanical connections, re-seat if needed, and run diagnostics if the issue persists.

    What NVIDIA says (3)

    “This event is logged when the GPU detects that a problem with a connection from the GPU to another GPU or NVSwitch over NVLink.”

    — NVIDIA Xid Catalog

    “For example, if a GPU fails, another GPU connected to it over NVLink may report an Xid 74 simply because the link went down as a result.”

    — NVIDIA Xid Catalog

    “Check link mechanical connections and re-seat if a field resolution is required. Run diags if issue persists.”

    — NVIDIA Xid Catalog

  4. Xid 13 is logged for general user application faults. NVIDIA says it is typically an out-of-bounds error, where code walks past the end of an array. Hardware failures can show up as Xid 13, but NVIDIA calls that rare. Because only one person's code triggers it, start by debugging that code.

    What NVIDIA says (2)

    “This event is logged for general user application faults. Typically this is an out-of-bounds error where the user has walked past the end of an array, but could also be an illegal instruction, illegal register, or other case.”

    — NVIDIA Xid Catalog

    “In rare cases, it’s possible for a hardware failure or system software bugs to materialize as XID 13.”

    — NVIDIA Xid Catalog

  5. Xid 45 is logged when an application aborts and the kernel driver tears down its work on the GPU. NVIDIA lists Ctrl-C, GPU resets and SIGKILL as examples. In many cases it is not a bug but a user or system action. If other Xids appear with it, follow the guidance for those.

    What NVIDIA says (2)

    “This event is logged when the user application aborts and the kernel driver tears down the GPU application running on the GPU. Control-C, GPU resets, sigkill are all examples where the application is aborted and this event is created.”

    — NVIDIA Xid Catalog

    “In many cases, this is not indicative of a bug but rather a user or system action.”

    — NVIDIA Xid Catalog

  6. Row remapping is how Ampere-generation and newer GPUs retire a faulty row of memory by mapping it to a spare row. Xid 63 records a row-remapping event, and the catalog's immediate action is IGNORE. Xid 64 records a row-remapping failure, and the catalog says RESET_GPU, then CONTACT_SUPPORT. An RMA (return merchandise authorization) is the return-for-replacement process. NVIDIA says the RMA criteria are met when the row-remapping failure flag is set and validated by field diagnostics.

    What NVIDIA says (4)

    “63 INFOROM_DRAM_RETIREMENT_EVENT”

    — NVIDIA Xid Catalog

    “64 INFOROM_DRAM_RETIREMENT_FAILURE”

    — NVIDIA Xid Catalog

    “On GPUs that support row remapping, starting with NVIDIA® Ampere archtecture GPUs, these events provide details on row remapper activity.”

    — NVIDIA Xid Catalog

    “Regarding row-remapping failures, the RMA criteria is met when the row-remapping failure flag is set and validated by the field diagnostic.”

    — GPU Memory Error Management: RMA Policy Thresholds

  7. Error containment is an Ampere-generation feature. It limits the impact of an uncorrectable ECC error to the application that hit it. For Xid 94, NVIDIA says the error is contained to one application, which must be restarted. All other applications are unaffected. NVIDIA still recommends resetting the GPU when convenient.

    What NVIDIA says (2)

    “For Xid 94, these errors are contained to one application, and the application that encountered this error must be restarted. All other applications running at the time of the Xid are unaffected. It is recommended to reset the GPU when convenient.”

    — NVIDIA Xid Catalog

    “The benefit of error containment is being able to limit the impact of uncorrectable ECC errors on GPU applications.”

    — GPU Memory Error Management: Error Containment

  8. An uncontained error is one the GPU could not limit to a single application. For Xid 95, NVIDIA says the errors affect multiple applications and the GPU must be reset before applications can restart. So the operator drains the work from that GPU and resets it, instead of restarting just one service.

    What NVIDIA says (2)

    “For Xid 95, these errors affect multiple applications, and the affected GPU must be reset before applications can restart.”

    — NVIDIA Xid Catalog

    “95 ROBUST_CHANNEL_UNCONTAINED_ERROR”

    — NVIDIA Xid Catalog

Key terms

Try it

Sample question

dmesg on a GPU node shows 'NVRM: Xid (PCI:0000:83:00): 48'. What does Xid 48 mean, and what does NVIDIA's catalog say to do when it appears on its own?

Show the answer

Answer: A double-bit ECC error, which the GPU could not correct. Reset the GPU, then run field diagnostics.

An Xid is an error report that the NVIDIA driver writes to the kernel log. Xid 48 is a double-bit ECC error. ECC (error-correcting code) memory can fix single-bit errors, but not double-bit ones, so this error is uncorrectable. NVIDIA's Xid Catalog lists the recovery for Xid 48 seen alone as RESET_GPU and the follow-up as RUN_FIELDDIAG (run field diagnostics). If Xid 63 or 64 appears with it, the catalog says to drain the node and reset.

What NVIDIA says (3)

“48 ROBUST_CHANNEL_GPU_ECC_DBE Double Bit ECC Error”

— NVIDIA Xid Catalog

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— NVIDIA Xid Catalog

“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU w/ 63 or 64: DRAIN_AND_RESET Investagatory Action Solo: RUN_FIELDDIAG”

— NVIDIA Xid Catalog

Practice 3.1 (8 questions) Full AI Operations guide

← 2.10 DPUs in the data center · 3.2 Orchestration and job scheduling →