NCP-AIO glossary

The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.

A

Access Control Services (ACS)

A PCIe feature that can force device-to-device traffic up through the CPU, slowing GDS.

Objectives: 4.4

What NVIDIA says (1)

“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”

— GPUDirect Storage Best Practices Guide

B

Base Command Manager (BCM)

NVIDIA cluster management software that provisions nodes from software images, monitors them and installs workload managers.

Objectives: 1.1, 1.4, 4.3

What NVIDIA says (2)

“NVIDIA Mission Control leverages NVIDIA Base Command Manager (BCM) for foundational cluster-management tasks such as provisioning compute nodes, configuring software images, assigning roles, and general cluster administration.”

— NVIDIA Mission Control Administration Guide: Software Stack

“Provisioning : Centrally store and deploy OS images of the compute, management nodes, and other various services.”

— NVIDIA Mission Control Administration Guide: Overview

Base View

The BCM web interface for monitoring and managing the cluster in a browser.

Objectives: 1.2, 1.5

What NVIDIA says (2)

“Base View is the web application front end to cluster management in BCM.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“https://<host name or IP address>:8081/base-view”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Baseboard management controller (BMC)

A small computer on the server board for remote power, monitoring and console access over the out-of-band network.

Objectives: 1.6, 2.2

What NVIDIA says (2)

“Once the network has been created, all nodes must be assigned a BMC interface, of type bmc, on this network.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”

— NVIDIA Mission Control Administration Guide: Overview

C

Checkpoint

A saved copy of training progress so a stopped job can resume.

Objectives: 3.4

What NVIDIA says (1)

“Always use shared network storage (e.g., NFS). When a preempted workload is resumed, it may be scheduled on a different node than before.”

— NVIDIA Run:ai: Checkpointing Preemptible Training Workloads

cm-diagnose

The BCM utility, run on the head node, that gathers cluster data for support.

Objectives: 1.7

What NVIDIA says (1)

“The diagnostic utility cm-diagnose is run from the head node. It gathers data on the cluster that may help diagnose issues.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

CMDaemon

The BCM daemon on each node that carries out management commands and writes /var/log/cmdaemon.

Objectives: 4.3

What NVIDIA says (2)

“CMDaemon generates log messages in /var/log/cmdaemon from specific internal subsystems”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“A global debug mode can be enabled in CMDaemon using cmdaemonctl”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

cmsh

The BCM cluster management shell, a command line with modes such as device, category and user; changes take effect on commit.

Objectives: 1.5, 1.6

What NVIDIA says (2)

“Whenever any changes are made using cmsh , it is important to remember to commit them or else they will not go into effect.”

— NVIDIA Mission Control Administration Guide: Software Stack

“The interfaces submode is accessible from the device mode.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

D

Data Center GPU Manager (DCGM)

NVIDIA tools for GPU monitoring and diagnostics; dcgmi is its command line.

Objectives: 3.5

What NVIDIA says (2)

“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”

— DCGM Diagnostics

“You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”

— DCGM User Guide: Getting Started

Data processing unit (DPU)

A BlueField network card with its own Arm cores that can run infrastructure services apart from the host.

Objectives: 1.6, 1.11

What NVIDIA says (2)

“dpusettings ................... Enter DPU settings setup mode”

— NVIDIA Mission Control Administration Guide: Node and Category Management

“Kubelet automatically pulls the container image from NGC and spawns a pod that runs the container.”

— DOCA Container Deployment Guide

Deserved quota

The GPU share a Run:ai project is guaranteed; work beyond it is over quota and can be preempted.

Objectives: 3.6

What NVIDIA says (2)

“The inference workload is assigned to a project and is affected by the project’s quota.”

— NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”

— NVIDIA Run:ai: Over Quota, Fairness and Preemption

DGX SuperPOD

NVIDIA's reference data center design of DGX systems, fabrics, storage and management nodes.

Objectives: 2.2

What NVIDIA says (2)

“Storage traffic is dedicated to its own fabric to remove interference with the node-to-node application traffic that can degrade overall performance.”

— NVIDIA Mission Control Administration Guide: Overview

“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”

— NVIDIA Mission Control Administration Guide: Overview

DOCA service

A containerized NVIDIA service, pulled from NGC, that runs on the BlueField Arm cores.

Objectives: 1.11

What NVIDIA says (1)

“Kubelet automatically pulls the container image from NGC and spawns a pod that runs the container.”

— DOCA Container Deployment Guide

Drain

A Slurm node state that lets running jobs finish but accepts no new ones.

Objectives: 2.1

What NVIDIA says (2)

“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

“Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

E

Enroot

A tool that turns container images into unprivileged sandboxes; Pyxis uses it.

Objectives: 3.3

What NVIDIA says (1)

“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

etcd

The distributed key-value store that holds Kubernetes cluster state; it runs on an odd number of nodes.

Objectives: 1.10

What NVIDIA says (2)

“An etcd cluster—the Kubernetes distributed key-value storage—runs on an odd number (1, 3, 5 ...) of nodes.”

— NVIDIA Base Command Manager 11 Containerization Manual (PDF)

“a minimum of three nodes is recommended for etcd”

— NVIDIA Base Command Manager 11 Containerization Manual (PDF)

F

Fabric Manager (FM)

The service that configures NVSwitches so GPUs can talk over NVLink; its version must match the driver.

Objectives: 4.2

What NVIDIA says (2)

“To check FM, for Linux based OS distributions, run the following command: sudo systemctl status”

— NVIDIA Fabric Manager User Guide

“NVIDIA Fabric Manager Package (same version as the Driver package).”

— NVIDIA Fabric Manager User Guide

Fair share

Scheduling that balances priority by past use across accounts or projects.

Objectives: 3.6

What NVIDIA says (1)

“Similar to the Slurm command sshare, an administrator can display the Slurm account hierarchy with the fairshare command in cmsh”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

G

gdsio

The GDS benchmarking tool for storage read and write load.

Objectives: 4.5

What NVIDIA says (1)

“the number of processes/threads generating IO is critical to determining maximum performance. The gdsio tool provides for specifying random reads or random writes”

— GPUDirect Storage Configuration and Benchmarking Guide

GPU fractions

A Run:ai feature that splits one GPU memory and compute among several workloads.

Objectives: 3.6

What NVIDIA says (1)

“With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users”

— NVIDIA Run:ai: GPU Fractions

GPU time-slicing

Sharing a GPU between pods by taking turns, with no memory or fault isolation.

Objectives: 2.4

What NVIDIA says (2)

“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”

— Time-Slicing GPUs in Kubernetes

“The NVIDIA GPU Operator enables oversubscription of GPUs through a set of extended options for the NVIDIA Kubernetes Device Plugin .”

— Time-Slicing GPUs in Kubernetes

GPUDirect Storage (GDS)

A direct data path between storage and GPU memory that avoids a CPU bounce buffer.

Objectives: 4.4, 4.5

What NVIDIA says (2)

“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”

— GPUDirect Storage Best Practices Guide

“In the /etc/cufile.json file, verify that allow_compat_mode is set to true . gdscheck -p displays whether the allow_compat_mode property is set to true .”

— GPUDirect Storage Troubleshooting Guide

H

Head node

The BCM server that stores node images, provisions the other nodes and collects their metrics.

Objectives: 1.1

What NVIDIA says (2)

“Provisioning : Centrally store and deploy OS images of the compute, management nodes, and other various services.”

— NVIDIA Mission Control Administration Guide: Overview

“Metrics : System monitoring and reporting that gather all telemetry from each of the nodes.”

— NVIDIA Mission Control Administration Guide: Overview

Health check

A BCM test that marks a node fit or unfit; a node failing a prejob check is kept from running the job.

Objectives: 1.7, 2.1

What NVIDIA says (2)

“In cmsh, the statuses of the services are listed by running the latesthealthdata command (section 10.6.3) from device mode.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“A node that has failed a prejob health check is not allowed to run a job.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

High-speed storage (HSS)

Shared fast storage for all nodes, where datasets and checkpoints of running jobs belong.

Objectives: 2.2, 4.5

What NVIDIA says (1)

“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”

— NVIDIA Mission Control Administration Guide: Overview

Horizontal Pod Autoscaling (HPA)

Kubernetes adding or removing pod replicas based on metrics.

Objectives: 3.1

What NVIDIA says (1)

“Configuring Horizontal Pod Autoscaling # Prerequisites # Prometheus installed on your cluster.”

— NVIDIA NIM Operator: NIM Service

I

imageupdate

The cmsh command that syncs a running node with its image; it does a dry run unless you pass -w.

Objectives: 1.4

What NVIDIA says (1)

“Performing dry run (use synclog command to review result, then pass -w to perform real update)”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

IMEX daemon (IMEX)

The service that lets GPUs in an NVLink domain share memory across nodes; on GB200 it can run per job.

Objectives: 2.1

What NVIDIA says (1)

“Running the IMEX daemon per job has the advantage that one user running a job cannot read the memory of a job run by another user on another node.”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

IOMMU

The input-output memory management unit, which remaps device memory access and must be off for GDS.

Objectives: 4.5

What NVIDIA says (2)

“When the IOMMU setting is enabled, PCIe traffic will be routed through the CPU root ports.”

— GPUDirect Storage Best Practices Guide

“Before you install GDS, you must disable IOMMU.”

— GPUDirect Storage Best Practices Guide

K

Knative

A Kubernetes add-on for request-driven serving that Run:ai inference workloads need.

Objectives: 3.2

What NVIDIA says (1)

“Make sure Knative is properly installed by your administrator.”

— NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

Kubelet

The Kubernetes agent on a node; on a DPU it starts DOCA service pods from YAML files in /etc/kubelet.d.

Objectives: 1.11

What NVIDIA says (2)

“cp doca_firefly.yaml /etc/kubelet.d”

— DOCA Container Deployment Guide

“journalctl -u kubelet Examines the Kubelet logs. Useful when a pod/container fails to spawn.”

— BlueField Troubleshooting Guide: DOCA Services

L

LDAP

Lightweight Directory Access Protocol: the directory service BCM runs on its head nodes to store users and groups.

Objectives: 1.5

What NVIDIA says (2)

“Out of the box, BCM runs its own LDAP service to help manage users and groups.”

— NVIDIA Mission Control Administration Guide: Software Stack

“This centralized LDAP service runs on the head nodes of the BCM managed cluster.”

— NVIDIA Mission Control Administration Guide: Software Stack

M

Magnum IO

NVIDIA's family of data-movement libraries, including NCCL and GPUDirect Storage.

Objectives: 4.4

What NVIDIA says (2)

“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”

— NCCL Environment Variables

“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”

— GPUDirect Storage Best Practices Guide

Multi-Instance GPU (MIG)

A feature that splits one GPU into isolated GPU instances, each with its own memory and compute.

Objectives: 2.5

What NVIDIA says (2)

“single : MIG mode is enabled on all GPUs on a node.”

— GPU Operator with MIG

“$ nvidia-smi mig -lgip”

— MIG User Guide: Getting Started with MIG

Multi-Node NVLink (MNNVL)

NVLink extended across several systems through an NVLink Switch network.

Objectives: 2.2

What NVIDIA says (1)

“Multi-Node NVLink is a capability enabled over an NVLink Switch network where multiple systems are interconnected to form a large GPU memory fabric also known as an NVLink Domain.”

— NVIDIA Mission Control Administration Guide: Overview

N

NCCL

The NVIDIA Collective Communications Library that moves data between GPUs for multi-GPU jobs.

Objectives: 4.4

What NVIDIA says (2)

“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”

— NCCL Environment Variables

“The NCCL_P2P_DISABLE variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”

— NCCL Environment Variables

NGC

NVIDIA's catalog of GPU-optimized software; its container registry is nvcr.io.

Objectives: 3.7, 4.6

What NVIDIA says (2)

“nvcr.io : The name of the container registry, which for the NGC container registry is nvcr.io .”

— NGC Catalog User Guide

“If you choose not to add a tag to an image, by default the word “latest” is added as the tag, however all NGC containers have an explicit version tag.”

— NGC Catalog User Guide

NIM Operator

A Kubernetes operator that deploys NIM services and caches their models.

Objectives: 3.1

What NVIDIA says (2)

“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”

— NVIDIA NIM Operator

“One key benefit of using the NIM Operator is its ability to pre-cache models and datasets.”

— NVIDIA NIM Operator

Node category

A BCM group of nodes that share the same configuration.

Objectives: 1.3, 1.8

What NVIDIA says (2)

“A node category is a group of regular nodes that share the same configuration.”

— NVIDIA Mission Control Administration Guide: Node and Category Management

“Nodes are typically divided into categories based on hardware specifications and their specific purpose.”

— NVIDIA Mission Control Administration Guide: Node and Category Management

Node Feature Discovery (NFD)

A Kubernetes add-on that labels nodes with their hardware features; the GPU Operator can deploy it.

Objectives: 2.4

What NVIDIA says (1)

“If NFD is already running in the cluster, then you must disable deploying NFD when you install the Operator.”

— GPU Operator: Installing the NVIDIA GPU Operator

Node pool

A Run:ai set of nodes with its own scheduler instance and quotas.

Objectives: 2.3

What NVIDIA says (2)

“Once created, the new node pool is automatically assigned to all projects and departments with a quota of zero GPU resources, unlimited CPU resources, and over quota enabled”

— NVIDIA Run:ai: Node Pools

“Creating a new node pool creates a new instance of the NVIDIA Run:ai Scheduler .”

— NVIDIA Run:ai: Node Pools

NVIDIA Container Toolkit

Software whose runtime hook gives containers access to host GPUs.

Objectives: 4.1, 4.6

What NVIDIA says (2)

“Edit your runtime configuration under /etc/nvidia-container-runtime/config.toml and uncomment the debug=... line.”

— NVIDIA Container Toolkit: Troubleshooting

“When using the NVIDIA Container Runtime Hook (that is, the Docker --gpus flag or the NVIDIA Container Runtime in legacy mode) to inject requested GPUs and driver libraries into a container”

— NVIDIA Container Toolkit: Troubleshooting

NVIDIA GPU Operator

A Kubernetes operator that installs and manages the GPU driver, container toolkit, device plugin and related parts.

Objectives: 1.10, 2.4, 4.6

What NVIDIA says (2)

“The NVIDIA GPU Operator is not selected by default.”

— NVIDIA Base Command Manager 11 Containerization Manual (PDF)

“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”

— GPU Operator: Installing the NVIDIA GPU Operator

NVIDIA Mission Control

NVIDIA software that runs an AI factory, built on BCM for provisioning plus services for observability and automated recovery.

Objectives: 1.1

What NVIDIA says (2)

“NVIDIA Mission Control leverages NVIDIA Base Command Manager (BCM) for foundational cluster-management tasks such as provisioning compute nodes, configuring software images, assigning roles, and general cluster administration.”

— NVIDIA Mission Control Administration Guide: Software Stack

“Admin Kubernetes Nodes (x86) - x3: BCM-integrated infrastructure services including Observability Stack, Autonomous Hardware Recovery (AHR), and Autonomous Job Recovery (AJR)”

— NVIDIA Mission Control Administration Guide: Overview

NVIDIA NIM (NIM)

Containerized NVIDIA inference microservices that serve AI models.

Objectives: 3.1, 3.2

What NVIDIA says (2)

“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”

— NVIDIA NIM Operator

“The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”

— NVIDIA Run:ai Inference Overview

NVIDIA Run:ai

A Kubernetes platform that schedules AI workloads and shares GPUs between teams by quota.

Objectives: 1.12, 2.3

What NVIDIA says (2)

“The installation of NVIDIA Run:ai is done through the cm-kubernetes-setup tool included with BCM 11.”

— NVIDIA Mission Control Administration Guide: Run:ai Installation

“NVIDIA Run:ai uses Projects as the primary organization management unit.”

— NVIDIA Run:ai: Projects

nvidia-smi

The NVIDIA System Management Interface command for GPU status, monitoring and reset.

Objectives: 3.5

What NVIDIA says (2)

“This tool allows the user to see one line of monitoring data per monitoring cycle.”

— nvidia-smi documentation

“Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”

— nvidia-smi documentation

NVML

The NVIDIA Management Library, the API behind nvidia-smi; NVML errors in a container mean GPU access is broken.

Objectives: 4.1, 4.6

What NVIDIA says (2)

“On systems where systemd is used to manage the cgroups of the container, reloading the systemd unit files ( systemctl daemon-reload ) is sufficient to trigger container updates and cause a loss of GPU access.”

— NVIDIA Container Toolkit: Troubleshooting

“You need to delete the container after the issue occurs. When it is restarted, either manually or automatically depending on whether you use a container orchestration platform, it regains access to the GPU.”

— NVIDIA Container Toolkit: Troubleshooting

O

Out-of-band network (OOB)

A separate Ethernet network for BMC management, apart from job traffic.

Objectives: 2.2

What NVIDIA says (1)

“The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”

— NVIDIA Mission Control Administration Guide: Overview

P

Preemption

The scheduler pausing a lower-priority workload to free its GPUs for another.

Objectives: 3.4, 3.6

What NVIDIA says (2)

“By default, training workloads are assigned a Low priority and are preemptible .”

— NVIDIA Run:ai: Train Models Using a Standard Training Workload

“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”

— NVIDIA Run:ai: Over Quota, Fairness and Preemption

Project manager

A BCM user allowed to see the job data of other named users for reports.

Objectives: 1.9

What NVIDIA says (1)

“alice can be made the project manager of bob and charlie. This allows her access to the job data of her subordinates”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Pyxis

A Slurm plugin that runs job steps in containers via srun --container-image.

Objectives: 3.3

What NVIDIA says (2)

“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“srun --container-image=ubuntu grep PRETTY /etc/os-release pyxis: importing docker image: ubuntu”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

R

Redfish

A standard REST API for managing servers through their BMC, used by BCM 11 for BIOS and firmware.

Objectives: 1.4

What NVIDIA says (2)

“the modern way of managing BIOS and firmware is with the Redfish standard.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“Firmware updated via Redfish need not be just the PC main system BIOS, but can also be the flashable software of subsystems, for example: NICs.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Run:ai department

A group of Run:ai projects under one shared scope.

Objectives: 2.3

What NVIDIA says (1)

“Departments group multiple projects under a shared organizational scope.”

— NVIDIA Run:ai: Departments

Run:ai project

The main Run:ai unit for a team: it holds a GPU quota and maps to a Kubernetes namespace.

Objectives: 2.3, 3.2

What NVIDIA says (2)

“NVIDIA Run:ai uses Projects as the primary organization management unit.”

— NVIDIA Run:ai: Projects

“Projects are manifested as Kubernetes namespaces.”

— NVIDIA Run:ai: Projects

S

Slurm

An open-source workload manager that queues batch jobs and allocates nodes and GPUs to them.

Objectives: 1.13, 2.1, 3.3

What NVIDIA says (2)

“The recommended way to run the cm-wlm-setup utility is without options or arguments, in which case a TUI dialog starts up.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

Slurm partition

A Slurm job queue that groups nodes and sets limits for them.

Objectives: 2.1

What NVIDIA says (2)

“A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

“append allowaccounts <user_id>”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

Software image

The directory tree of a node OS kept on the head node and copied to nodes at provisioning.

Objectives: 1.4, 1.8

What NVIDIA says (2)

“root@basecm11:~# cm-chroot-sw-img /cm/images/default-image”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“categories can share software images. It is not a one-to-one mapping between categories and software images.”

— NVIDIA Mission Control Administration Guide: Node and Category Management

T

topology/block

A Slurm topology plugin that BCM 11 uses for NVLink-aware scheduling on GB200 racks.

Objectives: 1.13

What NVIDIA says (1)

“BCM 11 introduces support for topology/block. This is required for BCM to enable NVLINK-aware scheduling of Slurm jobs for GB200/GB300 systems.”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

Triton Inference Server (Triton)

NVIDIA's open-source inference server, shipped as an NGC container.

Objectives: 4.6

What NVIDIA says (1)

“The HTTP request returns status 200 if Triton is ready and non-200 if it is not ready.”

— Quickstart — NVIDIA Triton Inference Server

Trusted TLS certificate

A certificate signed by an authority clients trust; Run:ai does not accept self-signed ones.

Objectives: 1.12

What NVIDIA says (1)

“TLS Certificate must be trusted. Self-signed certificates are not supported.”

— NVIDIA Run:ai: Install Using Base Command Manager

W

Workload manager (WLM)

Software such as Slurm that queues jobs and places them on nodes; BCM adds one with cm-wlm-setup.

Objectives: 1.3, 1.13

What NVIDIA says (2)

“A WLM may however also be added and configured after BCM has been installed, by using cm-wlm-setup, which is part of the cm-setup package, or by using the Base View WLM wizard.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“The recommended way to run the cm-wlm-setup utility is without options or arguments, in which case a TUI dialog starts up.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)