NCP-AII glossary

The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.

A

Access Control Services (ACS)

A PCIe feature that can force peer-to-peer traffic up to the CPU, which slows GDS and GPUDirect.

Objectives: 5.4, 5.5

What NVIDIA says (2)

“You can check whether ACS is enabled on PCI bridges by running: sudo lspci -vvv | grep ACSCtl”

— NCCL Troubleshooting: GPU troubleshooting

“For optimal GDS performance, disable ACS.”

— GPUDirect Storage Best Practices Guide

B

Base Command Manager (BCM)

NVIDIA cluster management software that provisions nodes, manages categories and installs Slurm.

Objectives: 3.1, 3.2, 3.3

What NVIDIA says (2)

“License the cluster by running the request-license and providing the product key.”

— DGX SuperPOD Deployment Guide: Initial Cluster Setup

“After the installation is completed, using the bcm-validate-pod command, one can verify whether all the settings are applied correctly according to the specification.”

— DGX SuperPOD Deployment Guide: DGX SuperPOD Validation

Baseboard management controller (BMC)

A small computer on the server board that lets you power, monitor and configure the system over the network, even when the OS is down.

Objectives: 1.3, 1.7

What NVIDIA says (2)

“Provides status and readings for system sensors, such as SSD, PSUs, voltages, CPU temperatures, DIMM temperatures, and fan speeds.”

— DGX H100/H200 User Guide: Using the BMC

“Set the IP address source to static. $ sudo ipmitool lan set 1 ipsrc static”

— DGX H100/H200 User Guide: Using the BMC

BFB image (BFB)

The bundled BlueField software image that bfb-install writes to the card through RShim.

Objectives: 2.1

What NVIDIA says (1)

“host# sudo bfb-install --rshim rshim<N> --bfb <image_path.bfb>”

— DOCA: BF-Bundle Installation and Upgrade

Bit error rate (BER)

The share of bits received wrongly on a link; mlxlink shows it with the physical counters.

Objectives: 4.4

What NVIDIA says (1)

“-c |--show_counters Show Physical Counters and BER Info”

— NVIDIA Firmware Tools: mlxlink Utility

Blanking panel

A cover for empty rack units so hot exhaust air cannot loop back to the server intakes.

Objectives: 1.5

What NVIDIA says (1)

“unoccupied RU spaces in the rack should be covered with blanking panels”

— DGX SuperPOD Data Center Design (H100): Cooling

BlueField DPU (DPU)

A network card with its own Arm cores; in DPU mode those cores own the NIC, in NIC mode they are off.

Objectives: 2.1, 4.7

What NVIDIA says (2)

“In DPU Mode, the NIC resources and functionality are owned and controlled by the embedded Arm subsystem.”

— DOCA: BlueField Modes of Operation

“In NIC Mode, BlueField operates as a ConnectX network adapter for the external host. For BlueField-3, the Arm cores are inactive”

— DOCA: BlueField Modes of Operation

Burn-in

Running a heavy workload for a long time to expose weak parts before production.

Objectives: 4.11, 4.12, 4.13

What NVIDIA says (2)

“-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”

— NVIDIA nccl-tests

“However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”

— DGX SuperPOD Administration Guide: System Health Checks and Debugging

Bus bandwidth (busbw)

The nccl-tests figure that reflects hardware link use, so it compares across GPU counts.

Objectives: 4.10

What NVIDIA says (1)

“See the Performance page for explanation about numbers, and in particular the "busbw" column.”

— NVIDIA nccl-tests

C

Cable validation

Comparing the real cabling, read from managed switches, with the planned topology.

Objectives: 4.5

What NVIDIA says (2)

“Cable validation is the process of validating the actual cable deployment, against the expected topology (from the planning).”

— InfiniBand Cluster Bring-up Procedure: Cable Validation

“The validation can utilize managed switches only.”

— InfiniBand Cluster Bring-up Procedure: Cable Validation

ClusterKit

A multipurpose node assessment tool that tests bandwidth, latency and more across cluster nodes.

Objectives: 4.9

What NVIDIA says (2)

“ClusterKit is a multipurpose node assessment tool for high-performance clusters”

— HPC-X: ClusterKit

“Message bandwidths (BWs) less than 93% (by default) of the maximum”

— HPC-X: Running ClusterKit

Customer-replaceable unit (CRU)

A component the customer may replace on site with an NVIDIA-supplied part.

Objectives: 1.9, 5.3

What NVIDIA says (2)

“You can obtain the following components for replacement in your data center.”

— DGX H100/H200 Service Manual: Introduction

“When replacing a component, use only the replacement supplied to you by NVIDIA.”

— DGX H100/H200 Service Manual: Introduction

D

DCGM diagnostics

dcgmi diag tests GPU health at run levels 1 to 4; higher levels run longer and include the lower ones.

Objectives: 1.10, 4.1

What NVIDIA says (2)

“higher numbered tests include all beneath.”

— DCGM User Guide: DCGM Diagnostics

“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”

— DCGM User Guide: DCGM Diagnostics

Digital Diagnostic Monitoring (DDM)

Live readings from a cable module, such as temperature and optical power.

Objectives: 1.8

What NVIDIA says (1)

“--ddm Get cable Digital Diagnostic Monitoring information”

— NVIDIA Firmware Tools: mlxlink Utility

DOCA

NVIDIA's software framework and driver stack for BlueField and ConnectX; DOCA-Host installs on the host.

Objectives: 3.4

What NVIDIA says (2)

“host# sudo apt install -y doca-all”

— DOCA: DOCA-Host Installation and Upgrade

“Full uninstallation is required before installing DOCA-Host:”

— DOCA: DOCA-Host Installation and Upgrade

E

Electrostatic discharge (ESD)

A static spark that can damage electronics; an ESD strap grounds you while you touch components.

Objectives: 1.9

What NVIDIA says (1)

“Wear an ESD strap during any procedure that involves touching electronic components.”

— DGX H100/H200 Service Manual: Motherboard Tray Removal and Installation

Enroot

A tool that turns container images into unprivileged sandboxes with no daemon.

Objectives: 3.3

What NVIDIA says (2)

“A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.”

— NVIDIA/enroot

“Standalone (no daemon)”

— NVIDIA/enroot

F

Fabric Manager (FM)

The service that sets up NVSwitch-based NVLink so GPUs can talk peer to peer.

Objectives: 4.3

What NVIDIA says (2)

“registers the daemon as the nvidia-fabricmanager system service.”

— NVIDIA Fabric Manager User Guide

“To successfully restore GPU NVLink peer-to-peer capability after the MIG mode is disabled on these systems, the FM service must be running.”

— NVIDIA Fabric Manager User Guide

First boot setup

The wizard that runs the first time a DGX starts, creating the admin account and the primary network setup.

Objectives: 1.6

What NVIDIA says (2)

“Create an administrative user account for the system, BMC, and Grub boot loader.”

— DGX H100/H200 User Guide: First Boot Setup

“Configure the primary network interface.”

— DGX H100/H200 User Guide: First Boot Setup

G

gdsio

The GDS benchmarking tool; -x picks the data path, -w the worker count and -T the run time.

Objectives: 4.14

What NVIDIA says (2)

“-x 0 , the IO data path, in this case GDS.”

— GPUDirect Storage Configuration and Benchmarking Guide

“-w 8 , 8 workers (8 IO threads)”

— GPUDirect Storage Configuration and Benchmarking Guide

GPU peer-to-peer (P2P)

GPUs reading and writing each other memory directly over NVLink or PCIe.

Objectives: 4.3

What NVIDIA says (2)

“You can use nvidia-smi topo -p2p <capability> to print a matrix of P2P status between GPU pairs.”

— NCCL Troubleshooting: GPU troubleshooting

“The <capability> value is p for PCIe and n for NVLink.”

— NCCL Troubleshooting: GPU troubleshooting

GPUDirect Storage (GDS)

A direct data path between storage and GPU memory that bypasses a CPU bounce buffer.

Objectives: 1.11, 4.14, 5.5

What NVIDIA says (2)

“xfer_type : 0 - Storage -> GPU ( GDS ) 1 - Storage -> CPU 2 - Storage -> CPU -> GPU”

— GPUDirect Storage Configuration and Benchmarking Guide

“Local drive configurations (Direct Attached Storage - DAS) and Network storage (Network Attached Storage - NAS) are covered.”

— GPUDirect Storage Configuration and Benchmarking Guide

H

Head node high availability (HA)

A BCM setup with a primary and secondary head node, a shared virtual IP and failover.

Objectives: 3.1

What NVIDIA says (2)

“Start the cmha-setup CLI wizard as the root user on the primary head node.”

— DGX SuperPOD Deployment Guide: High Availability

“This will be the IP that should always be used for accessing the active head nodes.”

— DGX SuperPOD Deployment Guide: High Availability

High-Performance Linpack (HPL)

A math-heavy benchmark used to load and compare systems; NVIDIA HPL runs one GPU per MPI process.

Objectives: 4.2, 4.12

What NVIDIA says (2)

“The NVIDIA HPL benchmark expects one GPU per MPI process. As such, set the number of MPI processes to match the number of available GPUs in the cluster.”

— NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

“Math intensive applications with network communications”

— DGX SuperPOD Administration Guide: System Health Checks and Debugging

I

ibdiagnet

An InfiniBand tool that scans the whole fabric and reports connectivity, devices and link width and speed.

Objectives: 4.4, 4.5

What NVIDIA says (2)

“Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Link width and speed checks”

— MLNX_OFED: InfiniBand Fabric Utilities

In-band management network

The normal OS network of each node, which carries cluster services such as BCM, Slurm and access to NGC and the home filesystem.

Objectives: 1.2

What NVIDIA says (1)

“Provides connectivity for the in-cluster services such as Base Command Manager, Slurm and to other services outside of the cluster such as the NGC registry, code repositories, and data sources.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

Intelligent Platform Management Interface (IPMI)

A standard protocol for talking to a BMC; ipmitool uses it, and it runs over RMCP+ on port 623.

Objectives: 1.3

What NVIDIA says (2)

“443 Redfish Redfish https with auth 623 RMCP+ IPMI”

— DGX H100/H200 User Guide: Using the BMC

“Set the IP address source to static. $ sudo ipmitool lan set 1 ipsrc static”

— DGX H100/H200 User Guide: Using the BMC

IOMMU

The CPU unit that remaps device memory access; it can redirect GPU peer-to-peer traffic to the CPU.

Objectives: 5.4

What NVIDIA says (2)

“IO virtualization (also known as VT-d or IOMMU) can interfere with GPU Direct by redirecting all PCI point-to-point traffic to the CPU root complex, causing a significant performance reduction or even a hang.”

— NCCL Troubleshooting: GPU troubleshooting

“For AMD CPUs, add amd_iommu=off . For Intel CPUs, add intel_iommu=off .”

— NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

M

mlxfwmanager

The NVIDIA tool that queries and updates firmware on NVIDIA network adapters.

Objectives: 1.4

What NVIDIA says (2)

“The mlxfwmanager is a firmware update and query utility which scans the system for available NVIDIA devices (only mst PCI devices) and performs the necessary firmware updates.”

— NVIDIA Firmware Tools: mlxfwmanager

“To query all the devices on the machine, use the following command line: # mlxfwmanager --query”

— NVIDIA Firmware Tools: mlxfwmanager

The NVIDIA tool that checks link state, cable or module details, error counters and the signal eye on a port.

Objectives: 1.8, 4.4

What NVIDIA says (2)

“The mlxlink tool is used to check and debug link status and related issues.”

— NVIDIA Firmware Tools: mlxlink Utility

“-m |--show_module Show Module Info”

— NVIDIA Firmware Tools: mlxlink Utility

Multi-Instance GPU (MIG)

A feature that splits one GPU into isolated GPU instances, each with compute instances inside.

Objectives: 2.2

What NVIDIA says (2)

“By default, MIG mode is not enabled on the GPU.”

— MIG User Guide: Getting Started with MIG

“Once the GPU instances are created, you need to create the corresponding Compute Instances (CI).”

— MIG User Guide: Getting Started with MIG

N

N+1 power

Power provisioning with one more feed than the load needs, so one feed can fail without stopping the system.

Objectives: 1.5

What NVIDIA says (2)

“the data center must minimally provide N+1 power, where N equals two power sources. Each power source must be sized to support 50% of the total peak load.”

— DGX SuperPOD Data Center Design (H100): Electrical

“However, with the specified N+1 power provisioning, N equals two circuits.”

— DGX SuperPOD Data Center Design (H100): Cooling

nccl-tests

NVIDIA's programs that measure the speed and correctness of NCCL collectives such as all_reduce.

Objectives: 4.3, 4.10, 4.11

What NVIDIA says (2)

“These tests check both the performance and the correctness of NCCL operations.”

— NVIDIA nccl-tests

“See the Performance page for explanation about numbers, and in particular the "busbw" column.”

— NVIDIA nccl-tests

NGC

NVIDIA's catalog and registry of GPU software, reached with the NGC CLI and an API key.

Objectives: 3.7

What NVIDIA says (1)

“Here are some examples of using NGC API keys to authenticate with NGC CLI and Docker CLI”

— NGC Overview

Node category

A BCM group of nodes that share the same configuration and software image.

Objectives: 3.3

What NVIDIA says (1)

“the administrator creates a new category called misc. The default category default already exists in a newly installed cluster.”

— DGX SuperPOD Administration Guide: Provisioning Nodes

NUMA

Non-uniform memory access: each CPU socket has its own close memory, so processes should run near their GPU and NIC.

Objectives: 5.4, 5.5

What NVIDIA says (2)

“On NUMA systems, each rank should generally use CPU cores and host memory close to its GPU and, for multi-node jobs, its NIC.”

— NCCL Troubleshooting: Performance and tuning

“Use nvidia-smi topo -m and lscpu --extended=CPU,NODE,SOCKET,CORE to inspect GPU, NIC, CPU, and NUMA locality.”

— NCCL Troubleshooting: Performance and tuning

NVIDIA Container Toolkit

The software that lets containers use the host GPUs; nvidia-ctk configures Docker or containerd for it.

Objectives: 3.5, 3.6

What NVIDIA says (2)

“sudo nvidia-ctk runtime configure --runtime = docker”

— NVIDIA Container Toolkit: Installation Guide

“The nvidia-ctk command modifies the /etc/docker/daemon.json file on the host. The file is updated so that Docker can use the NVIDIA Container Runtime.”

— NVIDIA Container Toolkit: Installation Guide

NVIDIA System Management (NVSM)

The DGX tool that reports system health and runs stress tests from one command line.

Objectives: 1.1, 1.7, 4.1, 5.1

What NVIDIA says (2)

“NVIDIA provides customers a diagnostics and management tool called NVIDIA System Management, or NVSM.”

— DGX H100/H200 User Guide: Quick Start and Basic Operation

“To run the tests, use NVSM.”

— DGX H100/H200 Service Manual: Introduction

O

Out-of-band management network (OOB)

A separate network that connects only the management ports (BMCs, PDUs, switch management) of every device.

Objectives: 1.2, 1.3

What NVIDIA says (2)

“It connects the management ports of all devices including DGX and management servers, storage, networking gear, rack PDUs, and all other devices.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

“These are separate onto their own fabric because there is no use-case where users need access to these ports and are secured using logical network separation.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

P

Parameter-Set Identification (PSID)

A 16-character string in NIC firmware that must match the board, so the right image is flashed.

Objectives: 1.4

What NVIDIA says (1)

“PSID (Parameter-Set Identification) is a 16-ascii character string embedded in the firmware image”

— DGX H100/H200 Service Manual: Updating the ConnectX-7 Firmware

Power capping

Setting a maximum power draw for a GPU or system; the GPU applies the most conservative limit it receives.

Objectives: 1.5

What NVIDIA says (2)

“The GPU has three sources of power limits:”

— DGX H100/H200 User Guide: Managing Power Capping

“The GPU Performance Monitoring Unit (PMU) selects the most conservative policy to cap power”

— DGX H100/H200 User Guide: Managing Power Capping

PXE boot (PXE)

Booting a node from the network so the provisioning node can send it a software image.

Objectives: 3.2

What NVIDIA says (2)

“Configure the DGX systems to PXE boot by default.”

— DGX SuperPOD Deployment Guide: Initial Cluster Setup

“The action of transferring the software image to the nodes is called node provisioning and is done by special nodes called the provisioning nodes.”

— DGX SuperPOD Administration Guide: Provisioning Nodes

Pyxis

A Slurm plugin that lets users run containers through srun.

Objectives: 3.3

What NVIDIA says (1)

“Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”

— NVIDIA/pyxis: Container plugin for Slurm

R

Rack power distribution unit (rPDU)

The power strip in a rack that feeds each server, usually fed from three-phase power.

Objectives: 1.5

What NVIDIA says (2)

“The preferred power for high-density deployment patterns is 415 VAC, 32A, three-phase, N+1.”

— DGX SuperPOD Data Center Design (H100): Electrical

“It connects the management ports of all devices including DGX and management servers, storage, networking gear, rack PDUs, and all other devices.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

RAID 1

A disk setup that keeps two identical copies of a drive, so the OS survives one drive failing.

Objectives: 1.7

What NVIDIA says (2)

“During this time, running the nvsm show health command reports a warning that the RAID volume is re-syncing.”

— DGX H100/H200 User Guide: First Boot Setup

“The process can take an hour to complete.”

— DGX H100/H200 User Guide: First Boot Setup

Rail-optimized fabric

A compute network where GPU n of every node connects to the same leaf switch, so same-rail traffic is one hop away.

Objectives: 1.2

What NVIDIA says (2)

“Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

“Traffic between nodes, or between rails, traverses the spine layer.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

Redfish

DMTF's standard REST API for managing and monitoring a server through its BMC.

Objectives: 1.3, 1.5

What NVIDIA says (2)

“Redfish is DMTF’s standard set of APIs for managing and monitoring a platform.”

— DGX H100/H200 User Guide: Redfish APIs Support

“By default, Redfish support is enabled in the DGX H100/H200 BMC and the SBIOS.”

— DGX H100/H200 User Guide: Redfish APIs Support

Return merchandise authorization (RMA)

The number NVIDIA support gives you to return a faulty part for repair or replacement.

Objectives: 5.3

What NVIDIA says (1)

“Contact NVIDIA Enterprise Support to obtain an RMA number for any system or component that needs to be returned for repair or replacement.”

— DGX H100/H200 Service Manual: Introduction

RShim

The host-side interface used to reach and install a BlueField device from the host.

Objectives: 2.1, 4.7

What NVIDIA says (1)

“To list the RShim devices present on your system, run the following command:”

— DOCA: BF-Bundle Installation and Upgrade

S

Scalable unit (SU)

The repeatable building block of a DGX SuperPOD: a fixed group of DGX systems with its own leaf switches.

Objectives: 1.2

What NVIDIA says (1)

“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems”

— DGX SuperPOD H100 Reference Architecture: Abstract

Secure Flash

A DGX protection that refuses firmware images that are not signed and verified.

Objectives: 1.4

What NVIDIA says (1)

“Secure Flash is implemented for the DGX H100/H200 to prevent unsigned and unverified firmware images from being flashed onto the system.”

— DGX H100/H200 User Guide: Security

Slurm

The workload manager that schedules jobs onto nodes and GPUs; BCM installs it with bcm-install-slurm.

Objectives: 3.3

What NVIDIA says (2)

“Run the bcm-install-slurm script.”

— DGX SuperPOD Deployment Guide: Slurm Setup

“For DGX H100 systems, generic resources are set to autodetect.”

— DGX SuperPOD Deployment Guide: Slurm Setup

System BIOS (SBIOS)

The firmware that starts the server before the OS loads; it holds settings such as TPM and boot order.

Objectives: 1.3, 3.2

What NVIDIA says (2)

“Here are some occasions where it might be necessary to reconfigure settings in the SBIOS:”

— DGX H100/H200 User Guide: SBIOS Settings

“Enabling the TPM and Preventing the BIOS from Sending Block SID Requests”

— DGX H100/H200 User Guide: SBIOS Settings

T

Transceiver

The module at a cable end that sends and receives the signal; optical and copper types use different firmware.

Objectives: 1.8, 4.8

What NVIDIA says (2)

“Note that we can have both optical & copper type transceivers - these can be identified by Vendor PN and running show explicitly.”

— InfiniBand Cluster Bring-up Procedure: Transceiver Firmware Installation

“each transceiver type has different firmware image.”

— InfiniBand Cluster Bring-up Procedure: Transceiver Firmware Installation

Trusted Platform Module (TPM)

A security chip that stores keys and measurements so the platform can prove it booted trusted code.

Objectives: 1.3

What NVIDIA says (1)

“Enabling the TPM and Preventing the BIOS from Sending Block SID Requests”

— DGX H100/H200 User Guide: SBIOS Settings

U

Unified Fabric Manager (UFM)

NVIDIA's management platform for InfiniBand fabrics, including firmware checks and a fabric health report.

Objectives: 4.6

What NVIDIA says (2)

“The high-speed InfiniBand fabrics are managed with NVIDIA Unified Fabric Manager (UFM).”

— DGX SuperPOD Administration Guide: Managing High-Speed Fabrics

“UFM fabric health report contains the results of a series of checks that run on the fabric.”

— InfiniBand Cluster Bring-up Procedure: UFM Fabric Health

X

Xid error

A numbered GPU driver error in the kernel log that points to a class of fault.

Objectives: 5.1

What NVIDIA says (2)

“This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”

— Xid Errors: Analyzing the Xid Catalog

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— Xid Errors: Analyzing the Xid Catalog