Control Plane Installation and Configuration
19% of the NCP-AII exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
System and Server Bring-up · Physical Layer Management · Control Plane Installation and Configuration · Cluster Test and Verification · Troubleshoot and Optimize
3.1 BCM install and HA
Installing Base Command Manager, licensing it and setting up a failover head node.
Key points
BCM (Base Command Manager) is NVIDIA's cluster manager. It provisions nodes and runs the cluster services. A head node is the server that runs BCM. HA means a second head node can take over if the first one fails. The guide starts HA with the cmha-setup wizard, run as root on the primary head node. The cluster nodes must be powered off first.
What NVIDIA says (2)
“Start the cmha-setup CLI wizard as the root user on the primary head node.”
“The cluster nodes must be powered off before configuring HA.”
In BCM HA, the secondary head node starts as a copy of the primary. You PXE boot it, choose RESCUE, and run /cm/cm-clone-install --failover. PXE (Preboot Execution Environment) means booting from the network instead of a local disk. Later, the Finalize step copies the MySQL database, which holds the cluster configuration.
What NVIDIA says (2)
“After the secondary head node has booted into the rescue environment, run the /cm/cm-clone-install --failover command, then enter YES when prompted.”
“This will clone the MySQL database from the primary to the secondary head node.”
A virtual IP (VIP) is an address that always points at whichever head node is active. Pinging it shows only that one head node answers. cmha status checks failover ping, MySQL and status from both sides. The active head node has an asterisk.
What NVIDIA says (2)
“The command tests the configuration from both directions: from the primary head node to the secondary, and from the secondary to the primary. The active head node is indicated by an asterisk.”
“This will be the IP that should always be used for accessing the active head nodes.”
NAS (network-attached storage) is a file server on the network. Both head nodes must see the same /cm/shared and /home, or a failover would lose files. So cmha-setup copies these directories to the NAS and mounts them everywhere.
What NVIDIA says (2)
“cmha-setup will copy the /cm/shared and /home directories to the shared storage and configure both head nodes and all cluster nodes to mount it.”
“must be stored on an NFS filesystem for HA availability.”
An ISO is a disk image file. The BCM installer ISO can be written to USB or mounted as virtual media through the BMC. After the install finishes and the head node reboots, the guide licenses the cluster with request-license.
What NVIDIA says (2)
“License the cluster by running the request-license and providing the product key.”
“Ensure that the BIOS of the target head node is configured in UEFI mode and that its boot order is configured to boot the media containing the BCM installer image.”
Key terms: Base Command Manager Head node high availability
3.2 Installing the OS on nodes
PXE boot and how BCM provisioning nodes push software images.
Key points
Provisioning means BCM sends a full software image to a node over the network. For that to work, the node must PXE boot, which means boot from the network. So the guide sets Boot Option #1 to [NETWORK] in the DGX BIOS.
What NVIDIA says (2)
“Configure the DGX systems to PXE boot by default.”
“enter the BIOS menu, and configure Boot Option #1 to be [NETWORK]”
A software image is the full OS file tree that a node runs. Provisioning copies that image to the node. A provisioning node is any node with the provisioning role. The head node always has it.
What NVIDIA says (2)
“The action of transferring the software image to the nodes is called node provisioning and is done by special nodes called the provisioning nodes.”
“Creating provisioning nodes is done by assigning a provisioning role to a node or category of nodes.”
Each provisioning node sends images to a limited number of nodes at the same time. That limit is provisioning slots. The default of 10 is safe for typical clusters. Setting it lower can prevent network and disk overload.
What NVIDIA says (2)
“The maximum number of nodes that can be provisioned in parallel by the provisioning node.”
“The default value is 10, which is safe for typical cluster setups.”
Key terms: System BIOS Base Command Manager PXE boot
3.3 Cluster setup: categories, Slurm, Enroot, Pyxis
Grouping nodes into categories and installing Slurm with GPU autodetect and container support.
Key points
cmsh is the BCM command-line shell. A category is a group of nodes that share settings, such as the software image and roles. In the guide, DGX H100 nodes are in category dgx-h100. You list nodes and categories with device list.
What NVIDIA says (3)
“Check the nodes and their categories.”
“device list -f hostname:20,category:10”
“the administrator creates a new category called misc. The default category default already exists in a newly installed cluster.”
Slurm is the workload manager. It queues jobs and assigns nodes and GPUs to them. On a SuperPOD, BCM installs it with the bcm-install-slurm script. Use -A for air-gapped mode, which means a site with no internet access.
What NVIDIA says (2)
“Run the bcm-install-slurm script.”
“Use the -A parameter to run the script in air-gapped mode.”
GRES (generic resources) is how Slurm tracks GPUs. gres.conf lists them. NVML (NVIDIA Management Library) is the driver library that reports GPU details. For DGX H100, the guide sets gpuautodetect to nvml. BCM then writes gres.conf for you.
What NVIDIA says (3)
“For DGX H100 systems, generic resources are set to autodetect.”
“set gpuautodetect nvml”
“The gres.conf file will be updated automatically by BCM”
SPANK is Slurm's plugin interface. Pyxis is NVIDIA's SPANK plugin. It lets unprivileged users run containers through srun, for example srun --container-image=almalinux:9.
What NVIDIA says (2)
“Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”
“srun --container-image=almalinux:9 grep PRETTY /etc/os-release”
A daemon is a background service. Docker uses one. Enroot does not. Enroot imports an image, creates a squashfs file from it, and starts it as an ordinary user. A squashfs is a compressed, read-only file system image.
What NVIDIA says (3)
“A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.”
“$ enroot import docker://ubuntu $ enroot create ubuntu.sqsh $ enroot start ubuntu”
“Standalone (no daemon)”
Key terms: Base Command Manager Node category Slurm Pyxis Enroot
Try it: Slurm Scheduler
3.4 GPU and DOCA drivers
Installing, checking, upgrading and removing the NVIDIA driver and DOCA-Host.
Key points
The driver includes kernel modules, which are code that runs inside the Linux kernel. They are built against kernel headers. If the headers do not match the running kernel, the build fails. So install the headers for uname -r before the driver.
What NVIDIA says (2)
“it is best to manually ensure the correct version of the kernel headers and development packages are installed prior to installing the NVIDIA driver”
“# apt install linux-headers-$(uname -r)”
The open kernel modules are NVIDIA's open-source GPU kernel driver. After enabling NVIDIA's repository, the guide installs them with apt install nvidia-open. For a compute-only server, it can install libnvidia-compute and nvidia-dkms-open instead.
What NVIDIA says (2)
“# apt install nvidia-open”
“apt -V install libnvidia-compute nvidia-dkms-open”
Installed is not the same as loaded. The kernel reports the loaded driver in /proc. The guide reads /proc/driver/nvidia/version.
What NVIDIA says (1)
“When the driver is loaded, the driver version can be found by executing the following command: $ cat /proc/driver/nvidia/version”
APT is Ubuntu's package manager. Pinning tells APT to stay on one driver branch or version. NVIDIA says the upgrade command stays the same. The pinning configuration decides where it lands.
What NVIDIA says (2)
“When upgrading the driver, whether configured to a pinned branch or the latest available, the command to execute is always the same; what matters is the APT pinning configuration:”
“# apt dist-upgrade”
DOCA-Host is NVIDIA's host software for ConnectX and BlueField: drivers, libraries and tools. OFED is the older InfiniBand and RDMA driver stack it replaced. Upgrading from 2.5.x needs a full uninstall first. From 2.6.0 or later, apt can upgrade in place.
What NVIDIA says (2)
“To upgrade from DOCA version 2.5.x, all DOCA and OFED related packages should be removed.”
“Full uninstallation is required before installing DOCA-Host:”
openibd is the service that loads the InfiniBand and RDMA drivers. MST (Mellanox Software Tools) creates device files used by the firmware tools. The DOCA guide updates firmware, loads the drivers, and initializes MST.
What NVIDIA says (4)
“host# sudo apt install -y doca-all”
“Update the firmware: host# sudo apt install -y mlnx-fw-updater”
“Load the drivers: host# sudo /etc/init.d/openibd restart”
“Initialize MST: host# sudo mst restart”
Removing packages with apt keeps the package database correct. Deleting files by hand does not. The guide purges a list of driver packages. The list includes nvidia-fabricmanager and nvidia-persistenced.
What NVIDIA says (2)
“Follow the below steps to properly uninstall the NVIDIA driver from your system.”
“# apt remove --autoremove --purge -V”
Key terms: DOCA
Try it: CUDA Stack Verification
3.5 NVIDIA Container Toolkit
Installing the toolkit and pointing Docker or containerd at the NVIDIA runtime.
Key points
The NVIDIA Container Toolkit lets containers use NVIDIA GPUs. 'nvidia-ctk runtime configure' updates /etc/docker/daemon.json so Docker can use the NVIDIA Container Runtime. Then you restart the Docker daemon.
What NVIDIA says (3)
“sudo nvidia-ctk runtime configure --runtime = docker”
“The nvidia-ctk command modifies the /etc/docker/daemon.json file on the host. The file is updated so that Docker can use the NVIDIA Container Runtime.”
“Restart the Docker daemon: $ sudo systemctl restart docker”
The NVIDIA Container Toolkit lets containers use the host's GPUs. It relies on the host driver. NVIDIA lists installing the driver as a prerequisite and recommends the package manager.
What NVIDIA says (2)
“Install the NVIDIA GPU driver for your Linux distribution.”
“NVIDIA recommends installing the driver by using the package manager for your distribution.”
Key terms: NVIDIA Container Toolkit
3.6 GPUs with Docker
Running a sample GPU container to prove the whole chain works.
Key points
A sample workload proves the whole chain works: driver, toolkit and Docker. NVIDIA runs nvidia-smi inside an ubuntu container with '--gpus all'. If the container prints the GPU table, Docker can use the GPUs.
What NVIDIA says (1)
“sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”
A sample workload proves the whole chain works: driver, toolkit, runtime and Docker. If nvidia-smi inside the container lists the GPUs, the setup is right.
What NVIDIA says (2)
“you can verify your installation by running a sample workload.”
“sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”
Key terms: NVIDIA Container Toolkit
Try it: NGC Container Flow
3.7 NGC CLI on hosts
Installing the NGC command line tool and authenticating it with an API key.
Key points
The NGC CLI is NVIDIA's command-line tool for the NGC catalog of containers and models. NVIDIA's examples authenticate it with an NGC API key: run 'ngc config set' and paste the key at the API_KEY prompt.
What NVIDIA says (2)
“Here are some examples of using NGC API keys to authenticate with NGC CLI and Docker CLI”
“ngc config set Paste your key value at the API_KEY prompt”
The NGC CLI is NVIDIA's command-line tool for the NGC catalog of containers and models. NVIDIA recommends always running the latest version for new features, bug fixes and security updates. The CLI can list its own releases and upgrade itself; the NGC CLI Installers page also has the downloads.
What NVIDIA says (2)
“Always use the latest NGC CLI version to access the newest features, bug fixes, performance improvements, and security updates.”
“or run ngc version list to view the latest releases, then upgrade using the following command: ngc version upgrade”
Key terms: NGC