AI Operations
22% of the NCA-AIIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Essential AI Knowledge · AI Infrastructure · AI Operations
3.1 Data center management and monitoring
Reading Xid errors and choosing the recovery action NVIDIA documents.
Key points
An Xid is an error report that the NVIDIA driver writes to the kernel log. Xid 48 is a double-bit ECC error. ECC (error-correcting code) memory can fix single-bit errors, but not double-bit ones, so this error is uncorrectable. NVIDIA's Xid Catalog lists the recovery for Xid 48 seen alone as RESET_GPU and the follow-up as RUN_FIELDDIAG (run field diagnostics). If Xid 63 or 64 appears with it, the catalog says to drain the node and reset.
What NVIDIA says (3)
“48 ROBUST_CHANNEL_GPU_ECC_DBE Double Bit ECC Error”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU w/ 63 or 64: DRAIN_AND_RESET Investagatory Action Solo: RUN_FIELDDIAG”
Xid 79 means the GPU has fallen off the bus. The bus here is PCI Express (PCIe), the link between the GPU and the server's CPU. NVIDIA says the event is logged when the driver tries to reach the GPU over PCIe and finds it is not accessible. It is often caused by hardware failures on the PCIe link, so check system and kernel PCI event logs.
What NVIDIA says (3)
“79 ROBUST_CHANNEL_GPU_HAS_FALLEN_OFF_THE_BUS”
“This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”
“This event is often caused by hardware failures on the PCI Express link causing the GPU to be inaccessible due to the link being brought down.”
NVLink is NVIDIA's direct GPU-to-GPU link. An NVSwitch is a switch chip that connects many GPUs over NVLink. Xid 74 is logged when the GPU detects a problem with an NVLink connection to another GPU or an NVSwitch. The cause can be at the far end. If one GPU fails, a GPU linked to it may report Xid 74 simply because the link went down. For several error bits, NVIDIA's workflow says to check link mechanical connections, re-seat if needed, and run diagnostics if the issue persists.
What NVIDIA says (3)
“This event is logged when the GPU detects that a problem with a connection from the GPU to another GPU or NVSwitch over NVLink.”
“For example, if a GPU fails, another GPU connected to it over NVLink may report an Xid 74 simply because the link went down as a result.”
“Check link mechanical connections and re-seat if a field resolution is required. Run diags if issue persists.”
Xid 13 is logged for general user application faults. NVIDIA says it is typically an out-of-bounds error, where code walks past the end of an array. Hardware failures can show up as Xid 13, but NVIDIA calls that rare. Because only one person's code triggers it, start by debugging that code.
What NVIDIA says (2)
“This event is logged for general user application faults. Typically this is an out-of-bounds error where the user has walked past the end of an array, but could also be an illegal instruction, illegal register, or other case.”
“In rare cases, it’s possible for a hardware failure or system software bugs to materialize as XID 13.”
Xid 45 is logged when an application aborts and the kernel driver tears down its work on the GPU. NVIDIA lists Ctrl-C, GPU resets and SIGKILL as examples. In many cases it is not a bug but a user or system action. If other Xids appear with it, follow the guidance for those.
What NVIDIA says (2)
“This event is logged when the user application aborts and the kernel driver tears down the GPU application running on the GPU. Control-C, GPU resets, sigkill are all examples where the application is aborted and this event is created.”
“In many cases, this is not indicative of a bug but rather a user or system action.”
Row remapping is how Ampere-generation and newer GPUs retire a faulty row of memory by mapping it to a spare row. Xid 63 records a row-remapping event, and the catalog's immediate action is IGNORE. Xid 64 records a row-remapping failure, and the catalog says RESET_GPU, then CONTACT_SUPPORT. An RMA (return merchandise authorization) is the return-for-replacement process. NVIDIA says the RMA criteria are met when the row-remapping failure flag is set and validated by field diagnostics.
What NVIDIA says (4)
“63 INFOROM_DRAM_RETIREMENT_EVENT”
“64 INFOROM_DRAM_RETIREMENT_FAILURE”
“On GPUs that support row remapping, starting with NVIDIA® Ampere archtecture GPUs, these events provide details on row remapper activity.”
“Regarding row-remapping failures, the RMA criteria is met when the row-remapping failure flag is set and validated by the field diagnostic.”
Error containment is an Ampere-generation feature. It limits the impact of an uncorrectable ECC error to the application that hit it. For Xid 94, NVIDIA says the error is contained to one application, which must be restarted. All other applications are unaffected. NVIDIA still recommends resetting the GPU when convenient.
What NVIDIA says (2)
“For Xid 94, these errors are contained to one application, and the application that encountered this error must be restarted. All other applications running at the time of the Xid are unaffected. It is recommended to reset the GPU when convenient.”
“The benefit of error containment is being able to limit the impact of uncorrectable ECC errors on GPU applications.”
An uncontained error is one the GPU could not limit to a single application. For Xid 95, NVIDIA says the errors affect multiple applications and the GPU must be reset before applications can restart. So the operator drains the work from that GPU and resets it, instead of restarting just one service.
What NVIDIA says (2)
“For Xid 95, these errors affect multiple applications, and the affected GPU must be reset before applications can restart.”
“95 ROBUST_CHANNEL_UNCONTAINED_ERROR”
3.2 Orchestration and job scheduling
How Kubernetes, Slurm, Run:ai and Base Command Manager hand out GPUs.
Key points
A device plugin is the Kubernetes mechanism that advertises special hardware to the scheduler. With NVIDIA's device plugin deployed, a container requests GPUs with the nvidia.com/gpu resource type. The scheduler then places the pod only on a node with that many free GPUs.
What NVIDIA says (1)
“With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”
Orchestration is the automated placement and management of workloads on shared resources. NVIDIA Run:ai is a GPU orchestration and optimization platform that aims to maximize compute utilization for AI workloads. Administrators define projects, departments and user roles, and Run:ai enforces policies for fair resource distribution.
What NVIDIA says (2)
“Allows administrators to define Projects, Departments and user roles, enforcing policies for fair resource distribution.”
“NVIDIA Run:ai is a GPU orchestration and optimization platform that helps organizations maximize compute utilization for AI workloads.”
Gang scheduling, also called batch scheduling, treats a group of pods as one unit. The NVIDIA Run:ai Scheduler documentation describes it as a workload of many pods being either fully scheduled or fully pending. That stops half of a job from starting and sitting idle.
What NVIDIA says (1)
“Gang scheduling describes a scheduling principle in which a workload composed of multiple pods is either fully scheduled (all pods are scheduled and running) or fully pending (none of the pods are running).”
Base Command Manager (BCM) is NVIDIA's cluster management software. The DGX SuperPOD reference architecture says it automates provisioning and administration. NVIDIA's product page says it integrates with tools like Slurm or NVIDIA Run:ai. Slurm is a widely used job scheduler for high-performance computing.
What NVIDIA says (3)
“It automates provisioning and administration, and supports clusters into the thousands of nodes”
“NVIDIA Base Command Manager streamlines cluster provisioning, workload management, and infrastructure monitoring.”
“Integrate with tools, like Slurm or NVIDIA Run:ai , to support traditional HPC or AI and analytics workloads across bare-metal and containerized environments.”
Kubernetes is a container orchestration system. An Operator is a Kubernetes pattern for automating the management of a software component. NVIDIA says configuring GPU nodes by hand needs several components, such as drivers and container runtimes, and is error-prone. The GPU Operator automates all of them: the driver, the device plugin, the Container Toolkit, GPU Feature Discovery (GFD) labels and DCGM (Data Center GPU Manager) monitoring.
What NVIDIA says (2)
“The NVIDIA GPU Operator uses the operator framework within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU.”
“These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the NVIDIA Container Toolkit , automatic node labeling using GFD , DCGM based monitoring and others.”
Key terms: GPU Operator Container Kubernetes Slurm
Try it: Slurm Scheduler Kubernetes GPU Ops
3.3 Monitoring GPUs
The DCGM metrics that matter and how to export them to Prometheus.
Key points
Field 311 counts double-bit volatile ECC errors. A double-bit error cannot be corrected by ECC, so any non-zero value is significant. NVIDIA's Xid Catalog logs uncorrectable errors as Xid 48, with RESET_GPU and RUN_FIELDDIAG as the actions when it appears alone.
What NVIDIA says (3)
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”
When a GPU lowers its clocks to protect itself, the reason is recorded as a clock event. DCGM field 112 reports the current clock event reasons as a bitmask, so you can see whether heat, power or something else caused the drop. Field 100 shows the clock value, but not the reason.
What NVIDIA says (2)
“DCGM_FI_DEV_CLOCKS_EVENT_REASONS 112 Current clock event reasons (bitmask of DCGM_CLOCKS_EVENT_REASON_*)”
“DCGM_FI_DEV_SM_CLOCK 100 SM clock for the device.”
DCGM (Data Center GPU Manager) exposes GPU telemetry as numbered fields. Field 310, DCGM_FI_DEV_ECC_SBE_VOL_TOTAL, is the total of single-bit volatile ECC errors. Volatile means the count since the driver last loaded. Single-bit errors are corrected, so the trend matters more than any one value.
What NVIDIA says (1)
“DCGM_FI_DEV_ECC_SBE_VOL_TOTAL 310 Total single bit volatile ECC errors.”
Prometheus is an open-source monitoring system that scrapes metrics over HTTP. DCGM Exporter exposes GPU metrics at an HTTP /metrics endpoint for tools such as Prometheus. Its default listen address is :9400.
What NVIDIA says (2)
“DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”
“Address of listening http server. Default: “:9400””
DCGM (Data Center GPU Manager) is NVIDIA's suite for managing and monitoring data center GPUs. NVIDIA says it includes active health monitoring, diagnostics, system alerts and governance policies such as power and clock management. Teams can use it alone or integrate it into cluster management and monitoring tools.
What NVIDIA says (5)
“Active Health Checks (GPU subsystems)”
“GPU Diagnostics (Diagnostic Levels - 1, 2, 3)”
“It includes active health monitoring, comprehensive diagnostics, system alerts, and governance policies including power and clock management.”
“The core DCGM library can be run as a standalone process or be loaded by an agent as a shared library.”
“Infrastructure teams can use it standalone and in addition easily integrate it into cluster management tools, resource scheduling, and monitoring products from NVIDIA partners.”
A streaming multiprocessor (SM) is one of the many processing blocks inside a GPU. DCGM defines graphics engine activity as the fraction of time any portion of the graphics or compute engines were active. So it can be high while most SMs are idle. DCGM's SM activity metric is the fraction of time at least one warp was active, averaged over all SMs. A warp is a group of threads the SM runs together. That metric shows how much of the chip is really working.
What NVIDIA says (2)
“Graphics Engine Activity The fraction of time any portion of the graphics or compute engines were active.”
“SM Activity The fraction of time at least one warp was active on a multiprocessor, averaged over all multiprocessors.”
Key terms: ECC DCGM DCGM Exporter Prometheus
Try it: ECC Error Lifecycle DCGM Monitoring
3.4 Virtualizing GPUs
MIG, time-slicing and vGPU: how each shares a GPU and what isolation it gives.
Key points
NVIDIA says MIG lets supported GPUs be securely partitioned into up to seven separate GPU instances. MIG mode is not enabled by default. An administrator enables it, then creates instances from profiles such as 1g.10gb on H100. A profile name gives the instance's share of compute (1g = one compute slice) and its memory size (10gb).
What NVIDIA says (3)
“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”
“By default, MIG mode is not enabled on the GPU.”
“Table 10 GPU Instance Profiles on H100”
A virtual machine (VM) is a software-defined computer that runs on a hypervisor. NVIDIA vGPU software is installed on a physical GPU in a server and creates virtual GPUs that can be shared across multiple VMs. NVIDIA says this raises GPU utilization while keeping management and security centralized.
What NVIDIA says (2)
“NVIDIA virtual GPU software enables multiple virtual machines (VMs) to have simultaneous, direct access to a single physical GPU”
“Installed on a physical GPU in a cloud or enterprise data center server, NVIDIA vGPU software creates virtual GPUs that can be shared across multiple virtual machines, scaling GPU utilization and enhancing ROI while centralizing manageability and security for enterprise IT.”
Multi-Instance GPU (MIG) partitions one physical GPU into separate GPU instances. NVIDIA says each instance gets dedicated compute and memory resources. Each instance's processors have isolated paths through the whole memory system. MIG also provides fault isolation between clients such as VMs, containers or processes.
What NVIDIA says (2)
“The Multi-Instance GPU (MIG) User Guide explains how to partition supported NVIDIA GPUs into multiple isolated instances, each with dedicated compute and memory resources.”
“MIG can partition available GPU compute resources (including streaming multiprocessors or SMs, and GPU engines such as copy engines or decoders), to provide a defined quality of service (QoS) with fault isolation for different clients such as VMs, containers or processes.”
Time-slicing lets workloads on an oversubscribed GPU take turns. Oversubscribed means more workloads than GPUs. NVIDIA says that, unlike MIG, there is no memory or fault isolation between the shared replicas. Time-slicing trades MIG's isolation for the ability to share a GPU among more users.
What NVIDIA says (2)
“GPU time-slicing enables workloads that are scheduled on oversubscribed GPUs to interleave with one another.”
“Time-slicing trades the memory and fault-isolation that is provided by MIG for the ability to share a GPU by a larger number of users.”
Key terms: Multi-Instance GPU Time-slicing Virtual GPU Streaming multiprocessor
Try it: MIG Partitioning GPU Virtualization