NCP-AIO glossary
The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.
A
- Access Control Services
A PCIe feature that can force device-to-device traffic up through the CPU, slowing GDS.
What NVIDIA says (1)
“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”
B
- Base Command Manager
NVIDIA cluster management software that provisions nodes from software images, monitors them and installs workload managers.
What NVIDIA says (2)
“NVIDIA Mission Control leverages NVIDIA Base Command Manager (BCM) for foundational cluster-management tasks such as provisioning compute nodes, configuring software images, assigning roles, and general cluster administration.”
“Provisioning : Centrally store and deploy OS images of the compute, management nodes, and other various services.”
- Base View
The BCM web interface for monitoring and managing the cluster in a browser.
What NVIDIA says (2)
“Base View is the web application front end to cluster management in BCM.”
“https://<host name or IP address>:8081/base-view”
- Baseboard management controller
A small computer on the server board for remote power, monitoring and console access over the out-of-band network.
What NVIDIA says (2)
“Once the network has been created, all nodes must be assigned a BMC interface, of type bmc, on this network.”
“The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”
C
- Checkpoint
A saved copy of training progress so a stopped job can resume.
What NVIDIA says (1)
“Always use shared network storage (e.g., NFS). When a preempted workload is resumed, it may be scheduled on a different node than before.”
- cm-diagnose
The BCM utility, run on the head node, that gathers cluster data for support.
What NVIDIA says (1)
“The diagnostic utility cm-diagnose is run from the head node. It gathers data on the cluster that may help diagnose issues.”
- CMDaemon
The BCM daemon on each node that carries out management commands and writes /var/log/cmdaemon.
What NVIDIA says (2)
“CMDaemon generates log messages in /var/log/cmdaemon from specific internal subsystems”
“A global debug mode can be enabled in CMDaemon using cmdaemonctl”
- cmsh
The BCM cluster management shell, a command line with modes such as device, category and user; changes take effect on commit.
What NVIDIA says (2)
“Whenever any changes are made using cmsh , it is important to remember to commit them or else they will not go into effect.”
“The interfaces submode is accessible from the device mode.”
D
- Data Center GPU Manager
NVIDIA tools for GPU monitoring and diagnostics; dcgmi is its command line.
What NVIDIA says (2)
“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”
“You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”
- Data processing unit
A BlueField network card with its own Arm cores that can run infrastructure services apart from the host.
What NVIDIA says (2)
“dpusettings ................... Enter DPU settings setup mode”
“Kubelet automatically pulls the container image from NGC and spawns a pod that runs the container.”
- Deserved quota
The GPU share a Run:ai project is guaranteed; work beyond it is over quota and can be preempted.
What NVIDIA says (2)
“The inference workload is assigned to a project and is affected by the project’s quota.”
“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”
- DGX SuperPOD
NVIDIA's reference data center design of DGX systems, fabrics, storage and management nodes.
What NVIDIA says (2)
“Storage traffic is dedicated to its own fabric to remove interference with the node-to-node application traffic that can degrade overall performance.”
“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”
- DOCA service
A containerized NVIDIA service, pulled from NGC, that runs on the BlueField Arm cores.
What NVIDIA says (1)
“Kubelet automatically pulls the container image from NGC and spawns a pod that runs the container.”
- Drain
A Slurm node state that lets running jobs finish but accepts no new ones.
What NVIDIA says (2)
“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”
“Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”
E
- Enroot
A tool that turns container images into unprivileged sandboxes; Pyxis uses it.
What NVIDIA says (1)
“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”
- etcd
The distributed key-value store that holds Kubernetes cluster state; it runs on an odd number of nodes.
What NVIDIA says (2)
“An etcd cluster—the Kubernetes distributed key-value storage—runs on an odd number (1, 3, 5 ...) of nodes.”
“a minimum of three nodes is recommended for etcd”
F
- Fabric Manager
The service that configures NVSwitches so GPUs can talk over NVLink; its version must match the driver.
What NVIDIA says (2)
“To check FM, for Linux based OS distributions, run the following command: sudo systemctl status”
“NVIDIA Fabric Manager Package (same version as the Driver package).”
Scheduling that balances priority by past use across accounts or projects.
What NVIDIA says (1)
“Similar to the Slurm command sshare, an administrator can display the Slurm account hierarchy with the fairshare command in cmsh”
G
- gdsio
The GDS benchmarking tool for storage read and write load.
What NVIDIA says (1)
“the number of processes/threads generating IO is critical to determining maximum performance. The gdsio tool provides for specifying random reads or random writes”
- GPU fractions
A Run:ai feature that splits one GPU memory and compute among several workloads.
What NVIDIA says (1)
“With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users”
- GPU time-slicing
Sharing a GPU between pods by taking turns, with no memory or fault isolation.
What NVIDIA says (2)
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”
“The NVIDIA GPU Operator enables oversubscription of GPUs through a set of extended options for the NVIDIA Kubernetes Device Plugin .”
- GPUDirect Storage
A direct data path between storage and GPU memory that avoids a CPU bounce buffer.
What NVIDIA says (2)
“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”
“In the /etc/cufile.json file, verify that allow_compat_mode is set to true . gdscheck -p displays whether the allow_compat_mode property is set to true .”
H
- Head node
The BCM server that stores node images, provisions the other nodes and collects their metrics.
What NVIDIA says (2)
“Provisioning : Centrally store and deploy OS images of the compute, management nodes, and other various services.”
“Metrics : System monitoring and reporting that gather all telemetry from each of the nodes.”
- Health check
A BCM test that marks a node fit or unfit; a node failing a prejob check is kept from running the job.
What NVIDIA says (2)
“In cmsh, the statuses of the services are listed by running the latesthealthdata command (section 10.6.3) from device mode.”
“A node that has failed a prejob health check is not allowed to run a job.”
- High-speed storage
Shared fast storage for all nodes, where datasets and checkpoints of running jobs belong.
What NVIDIA says (1)
“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”
- Horizontal Pod Autoscaling
Kubernetes adding or removing pod replicas based on metrics.
What NVIDIA says (1)
“Configuring Horizontal Pod Autoscaling # Prerequisites # Prometheus installed on your cluster.”
I
- imageupdate
The cmsh command that syncs a running node with its image; it does a dry run unless you pass -w.
What NVIDIA says (1)
“Performing dry run (use synclog command to review result, then pass -w to perform real update)”
- IMEX daemon
The service that lets GPUs in an NVLink domain share memory across nodes; on GB200 it can run per job.
What NVIDIA says (1)
“Running the IMEX daemon per job has the advantage that one user running a job cannot read the memory of a job run by another user on another node.”
- IOMMU
The input-output memory management unit, which remaps device memory access and must be off for GDS.
What NVIDIA says (2)
“When the IOMMU setting is enabled, PCIe traffic will be routed through the CPU root ports.”
“Before you install GDS, you must disable IOMMU.”
K
- Knative
A Kubernetes add-on for request-driven serving that Run:ai inference workloads need.
What NVIDIA says (1)
“Make sure Knative is properly installed by your administrator.”
- Kubelet
The Kubernetes agent on a node; on a DPU it starts DOCA service pods from YAML files in /etc/kubelet.d.
What NVIDIA says (2)
“cp doca_firefly.yaml /etc/kubelet.d”
“journalctl -u kubelet Examines the Kubelet logs. Useful when a pod/container fails to spawn.”
L
- LDAP
Lightweight Directory Access Protocol: the directory service BCM runs on its head nodes to store users and groups.
What NVIDIA says (2)
“Out of the box, BCM runs its own LDAP service to help manage users and groups.”
“This centralized LDAP service runs on the head nodes of the BCM managed cluster.”
M
- Magnum IO
NVIDIA's family of data-movement libraries, including NCCL and GPUDirect Storage.
What NVIDIA says (2)
“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”
“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”
- Multi-Instance GPU
A feature that splits one GPU into isolated GPU instances, each with its own memory and compute.
What NVIDIA says (2)
“single : MIG mode is enabled on all GPUs on a node.”
“$ nvidia-smi mig -lgip”
- Multi-Node NVLink
NVLink extended across several systems through an NVLink Switch network.
What NVIDIA says (1)
“Multi-Node NVLink is a capability enabled over an NVLink Switch network where multiple systems are interconnected to form a large GPU memory fabric also known as an NVLink Domain.”
N
- NCCL
The NVIDIA Collective Communications Library that moves data between GPUs for multi-GPU jobs.
What NVIDIA says (2)
“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”
“The NCCL_P2P_DISABLE variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”
- NGC
NVIDIA's catalog of GPU-optimized software; its container registry is nvcr.io.
What NVIDIA says (2)
“nvcr.io : The name of the container registry, which for the NGC container registry is nvcr.io .”
“If you choose not to add a tag to an image, by default the word “latest” is added as the tag, however all NGC containers have an explicit version tag.”
- NIM Operator
A Kubernetes operator that deploys NIM services and caches their models.
What NVIDIA says (2)
“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”
“One key benefit of using the NIM Operator is its ability to pre-cache models and datasets.”
- Node category
A BCM group of nodes that share the same configuration.
What NVIDIA says (2)
“A node category is a group of regular nodes that share the same configuration.”
“Nodes are typically divided into categories based on hardware specifications and their specific purpose.”
- Node Feature Discovery
A Kubernetes add-on that labels nodes with their hardware features; the GPU Operator can deploy it.
What NVIDIA says (1)
“If NFD is already running in the cluster, then you must disable deploying NFD when you install the Operator.”
- Node pool
A Run:ai set of nodes with its own scheduler instance and quotas.
What NVIDIA says (2)
“Once created, the new node pool is automatically assigned to all projects and departments with a quota of zero GPU resources, unlimited CPU resources, and over quota enabled”
“Creating a new node pool creates a new instance of the NVIDIA Run:ai Scheduler .”
- NVIDIA Container Toolkit
Software whose runtime hook gives containers access to host GPUs.
What NVIDIA says (2)
“Edit your runtime configuration under /etc/nvidia-container-runtime/config.toml and uncomment the debug=... line.”
“When using the NVIDIA Container Runtime Hook (that is, the Docker --gpus flag or the NVIDIA Container Runtime in legacy mode) to inject requested GPUs and driver libraries into a container”
- NVIDIA GPU Operator
A Kubernetes operator that installs and manages the GPU driver, container toolkit, device plugin and related parts.
What NVIDIA says (2)
“The NVIDIA GPU Operator is not selected by default.”
“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”
- NVIDIA Mission Control
NVIDIA software that runs an AI factory, built on BCM for provisioning plus services for observability and automated recovery.
What NVIDIA says (2)
“NVIDIA Mission Control leverages NVIDIA Base Command Manager (BCM) for foundational cluster-management tasks such as provisioning compute nodes, configuring software images, assigning roles, and general cluster administration.”
“Admin Kubernetes Nodes (x86) - x3: BCM-integrated infrastructure services including Observability Stack, Autonomous Hardware Recovery (AHR), and Autonomous Job Recovery (AJR)”
- NVIDIA NIM
Containerized NVIDIA inference microservices that serve AI models.
What NVIDIA says (2)
“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”
“The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”
- NVIDIA Run:ai
A Kubernetes platform that schedules AI workloads and shares GPUs between teams by quota.
What NVIDIA says (2)
“The installation of NVIDIA Run:ai is done through the cm-kubernetes-setup tool included with BCM 11.”
“NVIDIA Run:ai uses Projects as the primary organization management unit.”
- nvidia-smi
The NVIDIA System Management Interface command for GPU status, monitoring and reset.
What NVIDIA says (2)
“This tool allows the user to see one line of monitoring data per monitoring cycle.”
“Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”
- NVML
The NVIDIA Management Library, the API behind nvidia-smi; NVML errors in a container mean GPU access is broken.
What NVIDIA says (2)
“On systems where systemd is used to manage the cgroups of the container, reloading the systemd unit files ( systemctl daemon-reload ) is sufficient to trigger container updates and cause a loss of GPU access.”
“You need to delete the container after the issue occurs. When it is restarted, either manually or automatically depending on whether you use a container orchestration platform, it regains access to the GPU.”
O
- Out-of-band network
A separate Ethernet network for BMC management, apart from job traffic.
What NVIDIA says (1)
“The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”
P
- Preemption
The scheduler pausing a lower-priority workload to free its GPUs for another.
What NVIDIA says (2)
“By default, training workloads are assigned a Low priority and are preemptible .”
“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”
- Project manager
A BCM user allowed to see the job data of other named users for reports.
What NVIDIA says (1)
“alice can be made the project manager of bob and charlie. This allows her access to the job data of her subordinates”
- Pyxis
A Slurm plugin that runs job steps in containers via srun --container-image.
What NVIDIA says (2)
“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”
“srun --container-image=ubuntu grep PRETTY /etc/os-release pyxis: importing docker image: ubuntu”
R
- Redfish
A standard REST API for managing servers through their BMC, used by BCM 11 for BIOS and firmware.
What NVIDIA says (2)
“the modern way of managing BIOS and firmware is with the Redfish standard.”
“Firmware updated via Redfish need not be just the PC main system BIOS, but can also be the flashable software of subsystems, for example: NICs.”
- Run:ai department
A group of Run:ai projects under one shared scope.
What NVIDIA says (1)
“Departments group multiple projects under a shared organizational scope.”
- Run:ai project
The main Run:ai unit for a team: it holds a GPU quota and maps to a Kubernetes namespace.
What NVIDIA says (2)
“NVIDIA Run:ai uses Projects as the primary organization management unit.”
“Projects are manifested as Kubernetes namespaces.”
S
- Slurm
An open-source workload manager that queues batch jobs and allocates nodes and GPUs to them.
What NVIDIA says (2)
“The recommended way to run the cm-wlm-setup utility is without options or arguments, in which case a TUI dialog starts up.”
“A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”
- Slurm partition
A Slurm job queue that groups nodes and sets limits for them.
What NVIDIA says (2)
“A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”
“append allowaccounts <user_id>”
- Software image
The directory tree of a node OS kept on the head node and copied to nodes at provisioning.
What NVIDIA says (2)
“root@basecm11:~# cm-chroot-sw-img /cm/images/default-image”
“categories can share software images. It is not a one-to-one mapping between categories and software images.”
T
- topology/block
A Slurm topology plugin that BCM 11 uses for NVLink-aware scheduling on GB200 racks.
What NVIDIA says (1)
“BCM 11 introduces support for topology/block. This is required for BCM to enable NVLINK-aware scheduling of Slurm jobs for GB200/GB300 systems.”
- Triton Inference Server
NVIDIA's open-source inference server, shipped as an NGC container.
What NVIDIA says (1)
“The HTTP request returns status 200 if Triton is ready and non-200 if it is not ready.”
- Trusted TLS certificate
A certificate signed by an authority clients trust; Run:ai does not accept self-signed ones.
What NVIDIA says (1)
“TLS Certificate must be trusted. Self-signed certificates are not supported.”
W
- Workload manager
Software such as Slurm that queues jobs and places them on nodes; BCM adds one with cm-wlm-setup.
What NVIDIA says (2)
“A WLM may however also be added and configured after BCM has been installed, by using cm-wlm-setup, which is part of the cm-setup package, or by using the Base View WLM wizard.”
“The recommended way to run the cm-wlm-setup utility is without options or arguments, in which case a TUI dialog starts up.”