NCP-AII glossary
The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.
A
- Access Control Services
A PCIe feature that can force peer-to-peer traffic up to the CPU, which slows GDS and GPUDirect.
What NVIDIA says (2)
“You can check whether ACS is enabled on PCI bridges by running: sudo lspci -vvv | grep ACSCtl”
“For optimal GDS performance, disable ACS.”
B
- Base Command Manager
NVIDIA cluster management software that provisions nodes, manages categories and installs Slurm.
What NVIDIA says (2)
“License the cluster by running the request-license and providing the product key.”
“After the installation is completed, using the bcm-validate-pod command, one can verify whether all the settings are applied correctly according to the specification.”
- Baseboard management controller
A small computer on the server board that lets you power, monitor and configure the system over the network, even when the OS is down.
What NVIDIA says (2)
“Provides status and readings for system sensors, such as SSD, PSUs, voltages, CPU temperatures, DIMM temperatures, and fan speeds.”
“Set the IP address source to static. $ sudo ipmitool lan set 1 ipsrc static”
- BFB image
The bundled BlueField software image that bfb-install writes to the card through RShim.
What NVIDIA says (1)
“host# sudo bfb-install --rshim rshim<N> --bfb <image_path.bfb>”
- Bit error rate
The share of bits received wrongly on a link; mlxlink shows it with the physical counters.
What NVIDIA says (1)
“-c |--show_counters Show Physical Counters and BER Info”
- Blanking panel
A cover for empty rack units so hot exhaust air cannot loop back to the server intakes.
What NVIDIA says (1)
“unoccupied RU spaces in the rack should be covered with blanking panels”
- BlueField DPU
A network card with its own Arm cores; in DPU mode those cores own the NIC, in NIC mode they are off.
What NVIDIA says (2)
“In DPU Mode, the NIC resources and functionality are owned and controlled by the embedded Arm subsystem.”
“In NIC Mode, BlueField operates as a ConnectX network adapter for the external host. For BlueField-3, the Arm cores are inactive”
- Burn-in
Running a heavy workload for a long time to expose weak parts before production.
What NVIDIA says (2)
“-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”
“However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”
- Bus bandwidth
The nccl-tests figure that reflects hardware link use, so it compares across GPU counts.
What NVIDIA says (1)
“See the Performance page for explanation about numbers, and in particular the "busbw" column.”
C
- Cable validation
Comparing the real cabling, read from managed switches, with the planned topology.
What NVIDIA says (2)
“Cable validation is the process of validating the actual cable deployment, against the expected topology (from the planning).”
“The validation can utilize managed switches only.”
- ClusterKit
A multipurpose node assessment tool that tests bandwidth, latency and more across cluster nodes.
What NVIDIA says (2)
“ClusterKit is a multipurpose node assessment tool for high-performance clusters”
“Message bandwidths (BWs) less than 93% (by default) of the maximum”
- Customer-replaceable unit
A component the customer may replace on site with an NVIDIA-supplied part.
What NVIDIA says (2)
“You can obtain the following components for replacement in your data center.”
“When replacing a component, use only the replacement supplied to you by NVIDIA.”
D
- DCGM diagnostics
dcgmi diag tests GPU health at run levels 1 to 4; higher levels run longer and include the lower ones.
What NVIDIA says (2)
“higher numbered tests include all beneath.”
“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”
- Digital Diagnostic Monitoring
Live readings from a cable module, such as temperature and optical power.
What NVIDIA says (1)
“--ddm Get cable Digital Diagnostic Monitoring information”
- DOCA
NVIDIA's software framework and driver stack for BlueField and ConnectX; DOCA-Host installs on the host.
What NVIDIA says (2)
“host# sudo apt install -y doca-all”
“Full uninstallation is required before installing DOCA-Host:”
E
- Electrostatic discharge
A static spark that can damage electronics; an ESD strap grounds you while you touch components.
What NVIDIA says (1)
“Wear an ESD strap during any procedure that involves touching electronic components.”
- Enroot
A tool that turns container images into unprivileged sandboxes with no daemon.
What NVIDIA says (2)
“A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.”
“Standalone (no daemon)”
F
- Fabric Manager
The service that sets up NVSwitch-based NVLink so GPUs can talk peer to peer.
What NVIDIA says (2)
“registers the daemon as the nvidia-fabricmanager system service.”
“To successfully restore GPU NVLink peer-to-peer capability after the MIG mode is disabled on these systems, the FM service must be running.”
- First boot setup
The wizard that runs the first time a DGX starts, creating the admin account and the primary network setup.
What NVIDIA says (2)
“Create an administrative user account for the system, BMC, and Grub boot loader.”
“Configure the primary network interface.”
G
- gdsio
The GDS benchmarking tool; -x picks the data path, -w the worker count and -T the run time.
What NVIDIA says (2)
“-x 0 , the IO data path, in this case GDS.”
“-w 8 , 8 workers (8 IO threads)”
- GPU peer-to-peer
GPUs reading and writing each other memory directly over NVLink or PCIe.
What NVIDIA says (2)
“You can use nvidia-smi topo -p2p <capability> to print a matrix of P2P status between GPU pairs.”
“The <capability> value is p for PCIe and n for NVLink.”
- GPUDirect Storage
A direct data path between storage and GPU memory that bypasses a CPU bounce buffer.
What NVIDIA says (2)
“xfer_type : 0 - Storage -> GPU ( GDS ) 1 - Storage -> CPU 2 - Storage -> CPU -> GPU”
“Local drive configurations (Direct Attached Storage - DAS) and Network storage (Network Attached Storage - NAS) are covered.”
H
- Head node high availability
A BCM setup with a primary and secondary head node, a shared virtual IP and failover.
What NVIDIA says (2)
“Start the cmha-setup CLI wizard as the root user on the primary head node.”
“This will be the IP that should always be used for accessing the active head nodes.”
- High-Performance Linpack
A math-heavy benchmark used to load and compare systems; NVIDIA HPL runs one GPU per MPI process.
What NVIDIA says (2)
“The NVIDIA HPL benchmark expects one GPU per MPI process. As such, set the number of MPI processes to match the number of available GPUs in the cluster.”
“Math intensive applications with network communications”
I
- ibdiagnet
An InfiniBand tool that scans the whole fabric and reports connectivity, devices and link width and speed.
What NVIDIA says (2)
“Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”
“Link width and speed checks”
- In-band management network
The normal OS network of each node, which carries cluster services such as BCM, Slurm and access to NGC and the home filesystem.
What NVIDIA says (1)
“Provides connectivity for the in-cluster services such as Base Command Manager, Slurm and to other services outside of the cluster such as the NGC registry, code repositories, and data sources.”
- Intelligent Platform Management Interface
A standard protocol for talking to a BMC; ipmitool uses it, and it runs over RMCP+ on port 623.
What NVIDIA says (2)
“443 Redfish Redfish https with auth 623 RMCP+ IPMI”
“Set the IP address source to static. $ sudo ipmitool lan set 1 ipsrc static”
- IOMMU
The CPU unit that remaps device memory access; it can redirect GPU peer-to-peer traffic to the CPU.
What NVIDIA says (2)
“IO virtualization (also known as VT-d or IOMMU) can interfere with GPU Direct by redirecting all PCI point-to-point traffic to the CPU root complex, causing a significant performance reduction or even a hang.”
“For AMD CPUs, add amd_iommu=off . For Intel CPUs, add intel_iommu=off .”
M
- mlxfwmanager
The NVIDIA tool that queries and updates firmware on NVIDIA network adapters.
What NVIDIA says (2)
“The mlxfwmanager is a firmware update and query utility which scans the system for available NVIDIA devices (only mst PCI devices) and performs the necessary firmware updates.”
“To query all the devices on the machine, use the following command line: # mlxfwmanager --query”
- mlxlink
The NVIDIA tool that checks link state, cable or module details, error counters and the signal eye on a port.
What NVIDIA says (2)
“The mlxlink tool is used to check and debug link status and related issues.”
“-m |--show_module Show Module Info”
- Multi-Instance GPU
A feature that splits one GPU into isolated GPU instances, each with compute instances inside.
What NVIDIA says (2)
“By default, MIG mode is not enabled on the GPU.”
“Once the GPU instances are created, you need to create the corresponding Compute Instances (CI).”
N
- N+1 power
Power provisioning with one more feed than the load needs, so one feed can fail without stopping the system.
What NVIDIA says (2)
“the data center must minimally provide N+1 power, where N equals two power sources. Each power source must be sized to support 50% of the total peak load.”
“However, with the specified N+1 power provisioning, N equals two circuits.”
- nccl-tests
NVIDIA's programs that measure the speed and correctness of NCCL collectives such as all_reduce.
What NVIDIA says (2)
“These tests check both the performance and the correctness of NCCL operations.”
“See the Performance page for explanation about numbers, and in particular the "busbw" column.”
- NGC
NVIDIA's catalog and registry of GPU software, reached with the NGC CLI and an API key.
What NVIDIA says (1)
“Here are some examples of using NGC API keys to authenticate with NGC CLI and Docker CLI”
- Node category
A BCM group of nodes that share the same configuration and software image.
What NVIDIA says (1)
“the administrator creates a new category called misc. The default category default already exists in a newly installed cluster.”
- NUMA
Non-uniform memory access: each CPU socket has its own close memory, so processes should run near their GPU and NIC.
What NVIDIA says (2)
“On NUMA systems, each rank should generally use CPU cores and host memory close to its GPU and, for multi-node jobs, its NIC.”
“Use nvidia-smi topo -m and lscpu --extended=CPU,NODE,SOCKET,CORE to inspect GPU, NIC, CPU, and NUMA locality.”
- NVIDIA Container Toolkit
The software that lets containers use the host GPUs; nvidia-ctk configures Docker or containerd for it.
What NVIDIA says (2)
“sudo nvidia-ctk runtime configure --runtime = docker”
“The nvidia-ctk command modifies the /etc/docker/daemon.json file on the host. The file is updated so that Docker can use the NVIDIA Container Runtime.”
- NVIDIA System Management
The DGX tool that reports system health and runs stress tests from one command line.
What NVIDIA says (2)
“NVIDIA provides customers a diagnostics and management tool called NVIDIA System Management, or NVSM.”
“To run the tests, use NVSM.”
O
- Out-of-band management network
A separate network that connects only the management ports (BMCs, PDUs, switch management) of every device.
What NVIDIA says (2)
“It connects the management ports of all devices including DGX and management servers, storage, networking gear, rack PDUs, and all other devices.”
“These are separate onto their own fabric because there is no use-case where users need access to these ports and are secured using logical network separation.”
P
- Parameter-Set Identification
A 16-character string in NIC firmware that must match the board, so the right image is flashed.
What NVIDIA says (1)
“PSID (Parameter-Set Identification) is a 16-ascii character string embedded in the firmware image”
- Power capping
Setting a maximum power draw for a GPU or system; the GPU applies the most conservative limit it receives.
What NVIDIA says (2)
“The GPU has three sources of power limits:”
“The GPU Performance Monitoring Unit (PMU) selects the most conservative policy to cap power”
- PXE boot
Booting a node from the network so the provisioning node can send it a software image.
What NVIDIA says (2)
“Configure the DGX systems to PXE boot by default.”
“The action of transferring the software image to the nodes is called node provisioning and is done by special nodes called the provisioning nodes.”
- Pyxis
A Slurm plugin that lets users run containers through srun.
What NVIDIA says (1)
“Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”
R
- Rack power distribution unit
The power strip in a rack that feeds each server, usually fed from three-phase power.
What NVIDIA says (2)
“The preferred power for high-density deployment patterns is 415 VAC, 32A, three-phase, N+1.”
“It connects the management ports of all devices including DGX and management servers, storage, networking gear, rack PDUs, and all other devices.”
- RAID 1
A disk setup that keeps two identical copies of a drive, so the OS survives one drive failing.
What NVIDIA says (2)
“During this time, running the nvsm show health command reports a warning that the RAID volume is re-syncing.”
“The process can take an hour to complete.”
- Rail-optimized fabric
A compute network where GPU n of every node connects to the same leaf switch, so same-rail traffic is one hop away.
What NVIDIA says (2)
“Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”
“Traffic between nodes, or between rails, traverses the spine layer.”
- Redfish
DMTF's standard REST API for managing and monitoring a server through its BMC.
What NVIDIA says (2)
“Redfish is DMTF’s standard set of APIs for managing and monitoring a platform.”
“By default, Redfish support is enabled in the DGX H100/H200 BMC and the SBIOS.”
- Return merchandise authorization
The number NVIDIA support gives you to return a faulty part for repair or replacement.
What NVIDIA says (1)
“Contact NVIDIA Enterprise Support to obtain an RMA number for any system or component that needs to be returned for repair or replacement.”
- RShim
The host-side interface used to reach and install a BlueField device from the host.
What NVIDIA says (1)
“To list the RShim devices present on your system, run the following command:”
S
- Scalable unit
The repeatable building block of a DGX SuperPOD: a fixed group of DGX systems with its own leaf switches.
What NVIDIA says (1)
“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems”
- Secure Flash
A DGX protection that refuses firmware images that are not signed and verified.
What NVIDIA says (1)
“Secure Flash is implemented for the DGX H100/H200 to prevent unsigned and unverified firmware images from being flashed onto the system.”
- Slurm
The workload manager that schedules jobs onto nodes and GPUs; BCM installs it with bcm-install-slurm.
What NVIDIA says (2)
“Run the bcm-install-slurm script.”
“For DGX H100 systems, generic resources are set to autodetect.”
- System BIOS
The firmware that starts the server before the OS loads; it holds settings such as TPM and boot order.
What NVIDIA says (2)
“Here are some occasions where it might be necessary to reconfigure settings in the SBIOS:”
“Enabling the TPM and Preventing the BIOS from Sending Block SID Requests”
T
- Transceiver
The module at a cable end that sends and receives the signal; optical and copper types use different firmware.
What NVIDIA says (2)
“Note that we can have both optical & copper type transceivers - these can be identified by Vendor PN and running show explicitly.”
“each transceiver type has different firmware image.”
- Trusted Platform Module
A security chip that stores keys and measurements so the platform can prove it booted trusted code.
What NVIDIA says (1)
“Enabling the TPM and Preventing the BIOS from Sending Block SID Requests”
U
- Unified Fabric Manager
NVIDIA's management platform for InfiniBand fabrics, including firmware checks and a fabric health report.
What NVIDIA says (2)
“The high-speed InfiniBand fabrics are managed with NVIDIA Unified Fabric Manager (UFM).”
“UFM fabric health report contains the results of a series of checks that run on the fabric.”
X
- Xid error
A numbered GPU driver error in the kernel log that points to a class of fault.
What NVIDIA says (2)
“This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”