Control Plane Installation and Configuration

19% of the NCP-AII exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

System and Server Bring-up · Physical Layer Management · Control Plane Installation and Configuration · Cluster Test and Verification · Troubleshoot and Optimize

3.1 BCM install and HA

Official objective: “Install Base Command Manager (BCM), configure and verify HA.”

Installing Base Command Manager, licensing it and setting up a failover head node.

Key points

  1. BCM (Base Command Manager) is NVIDIA's cluster manager. It provisions nodes and runs the cluster services. A head node is the server that runs BCM. HA means a second head node can take over if the first one fails. The guide starts HA with the cmha-setup wizard, run as root on the primary head node. The cluster nodes must be powered off first.

    What NVIDIA says (2)

    “Start the cmha-setup CLI wizard as the root user on the primary head node.”

    — DGX SuperPOD Deployment Guide: High Availability

    “The cluster nodes must be powered off before configuring HA.”

    — DGX SuperPOD Deployment Guide: High Availability

  2. In BCM HA, the secondary head node starts as a copy of the primary. You PXE boot it, choose RESCUE, and run /cm/cm-clone-install --failover. PXE (Preboot Execution Environment) means booting from the network instead of a local disk. Later, the Finalize step copies the MySQL database, which holds the cluster configuration.

    What NVIDIA says (2)

    “After the secondary head node has booted into the rescue environment, run the /cm/cm-clone-install --failover command, then enter YES when prompted.”

    — DGX SuperPOD Deployment Guide: High Availability

    “This will clone the MySQL database from the primary to the secondary head node.”

    — DGX SuperPOD Deployment Guide: High Availability

  3. A virtual IP (VIP) is an address that always points at whichever head node is active. Pinging it shows only that one head node answers. cmha status checks failover ping, MySQL and status from both sides. The active head node has an asterisk.

    What NVIDIA says (2)

    “The command tests the configuration from both directions: from the primary head node to the secondary, and from the secondary to the primary. The active head node is indicated by an asterisk.”

    — DGX SuperPOD Deployment Guide: High Availability

    “This will be the IP that should always be used for accessing the active head nodes.”

    — DGX SuperPOD Deployment Guide: High Availability

  4. NAS (network-attached storage) is a file server on the network. Both head nodes must see the same /cm/shared and /home, or a failover would lose files. So cmha-setup copies these directories to the NAS and mounts them everywhere.

    What NVIDIA says (2)

    “cmha-setup will copy the /cm/shared and /home directories to the shared storage and configure both head nodes and all cluster nodes to mount it.”

    — DGX SuperPOD Deployment Guide: High Availability

    “must be stored on an NFS filesystem for HA availability.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

  5. An ISO is a disk image file. The BCM installer ISO can be written to USB or mounted as virtual media through the BMC. After the install finishes and the head node reboots, the guide licenses the cluster with request-license.

    What NVIDIA says (2)

    “License the cluster by running the request-license and providing the product key.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “Ensure that the BIOS of the target head node is configured in UEFI mode and that its boot order is configured to boot the media containing the BCM installer image.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

Key terms: Base Command Manager Head node high availability

Practice 3.1 (5 questions) Objective page

3.2 Installing the OS on nodes

Official objective: “Install OS.”

PXE boot and how BCM provisioning nodes push software images.

Key points

  1. Provisioning means BCM sends a full software image to a node over the network. For that to work, the node must PXE boot, which means boot from the network. So the guide sets Boot Option #1 to [NETWORK] in the DGX BIOS.

    What NVIDIA says (2)

    “Configure the DGX systems to PXE boot by default.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “enter the BIOS menu, and configure Boot Option #1 to be [NETWORK]”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

  2. A software image is the full OS file tree that a node runs. Provisioning copies that image to the node. A provisioning node is any node with the provisioning role. The head node always has it.

    What NVIDIA says (2)

    “The action of transferring the software image to the nodes is called node provisioning and is done by special nodes called the provisioning nodes.”

    — DGX SuperPOD Administration Guide: Provisioning Nodes

    “Creating provisioning nodes is done by assigning a provisioning role to a node or category of nodes.”

    — DGX SuperPOD Administration Guide: Provisioning Nodes

  3. Each provisioning node sends images to a limited number of nodes at the same time. That limit is provisioning slots. The default of 10 is safe for typical clusters. Setting it lower can prevent network and disk overload.

    What NVIDIA says (2)

    “The maximum number of nodes that can be provisioned in parallel by the provisioning node.”

    — DGX SuperPOD Administration Guide: Provisioning Nodes

    “The default value is 10, which is safe for typical cluster setups.”

    — DGX SuperPOD Administration Guide: Provisioning Nodes

Key terms: System BIOS Base Command Manager PXE boot

Practice 3.2 (3 questions) Objective page

3.3 Cluster setup: categories, Slurm, Enroot, Pyxis

Official objective: “Install Cluster (configure category, configure interfaces, install Slurm/Enroot/Pyxis).”

Grouping nodes into categories and installing Slurm with GPU autodetect and container support.

Key points

  1. cmsh is the BCM command-line shell. A category is a group of nodes that share settings, such as the software image and roles. In the guide, DGX H100 nodes are in category dgx-h100. You list nodes and categories with device list.

    What NVIDIA says (3)

    “Check the nodes and their categories.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “device list -f hostname:20,category:10”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “the administrator creates a new category called misc. The default category default already exists in a newly installed cluster.”

    — DGX SuperPOD Administration Guide: Provisioning Nodes

  2. Slurm is the workload manager. It queues jobs and assigns nodes and GPUs to them. On a SuperPOD, BCM installs it with the bcm-install-slurm script. Use -A for air-gapped mode, which means a site with no internet access.

    What NVIDIA says (2)

    “Run the bcm-install-slurm script.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

    “Use the -A parameter to run the script in air-gapped mode.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

  3. GRES (generic resources) is how Slurm tracks GPUs. gres.conf lists them. NVML (NVIDIA Management Library) is the driver library that reports GPU details. For DGX H100, the guide sets gpuautodetect to nvml. BCM then writes gres.conf for you.

    What NVIDIA says (3)

    “For DGX H100 systems, generic resources are set to autodetect.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

    “set gpuautodetect nvml”

    — DGX SuperPOD Deployment Guide: Slurm Setup

    “The gres.conf file will be updated automatically by BCM”

    — DGX SuperPOD Deployment Guide: Slurm Setup

  4. SPANK is Slurm's plugin interface. Pyxis is NVIDIA's SPANK plugin. It lets unprivileged users run containers through srun, for example srun --container-image=almalinux:9.

    What NVIDIA says (2)

    “Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”

    — NVIDIA/pyxis: Container plugin for Slurm

    “srun --container-image=almalinux:9 grep PRETTY /etc/os-release”

    — NVIDIA/pyxis: Container plugin for Slurm

  5. A daemon is a background service. Docker uses one. Enroot does not. Enroot imports an image, creates a squashfs file from it, and starts it as an ordinary user. A squashfs is a compressed, read-only file system image.

    What NVIDIA says (3)

    “A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.”

    — NVIDIA/enroot

    “$ enroot import docker://ubuntu $ enroot create ubuntu.sqsh $ enroot start ubuntu”

    — NVIDIA/enroot

    “Standalone (no daemon)”

    — NVIDIA/enroot

Key terms: Base Command Manager Node category Slurm Pyxis Enroot

Try it: Slurm Scheduler

Practice 3.3 (5 questions) Objective page

3.4 GPU and DOCA drivers

Official objective: “Install/update/remove NVIDIA GPU and DOCA drivers.”

Installing, checking, upgrading and removing the NVIDIA driver and DOCA-Host.

Key points

  1. The driver includes kernel modules, which are code that runs inside the Linux kernel. They are built against kernel headers. If the headers do not match the running kernel, the build fails. So install the headers for uname -r before the driver.

    What NVIDIA says (2)

    “it is best to manually ensure the correct version of the kernel headers and development packages are installed prior to installing the NVIDIA driver”

    — NVIDIA Driver Installation Guide: Pre-installation Actions

    “# apt install linux-headers-$(uname -r)”

    — NVIDIA Driver Installation Guide: Ubuntu

  2. The open kernel modules are NVIDIA's open-source GPU kernel driver. After enabling NVIDIA's repository, the guide installs them with apt install nvidia-open. For a compute-only server, it can install libnvidia-compute and nvidia-dkms-open instead.

    What NVIDIA says (2)

    “# apt install nvidia-open”

    — NVIDIA Driver Installation Guide: Ubuntu

    “apt -V install libnvidia-compute nvidia-dkms-open”

    — NVIDIA Driver Installation Guide: Ubuntu

  3. Installed is not the same as loaded. The kernel reports the loaded driver in /proc. The guide reads /proc/driver/nvidia/version.

    What NVIDIA says (1)

    “When the driver is loaded, the driver version can be found by executing the following command: $ cat /proc/driver/nvidia/version”

    — NVIDIA Driver Installation Guide: Post-installation Actions

  4. APT is Ubuntu's package manager. Pinning tells APT to stay on one driver branch or version. NVIDIA says the upgrade command stays the same. The pinning configuration decides where it lands.

    What NVIDIA says (2)

    “When upgrading the driver, whether configured to a pinned branch or the latest available, the command to execute is always the same; what matters is the APT pinning configuration:”

    — NVIDIA Driver Installation Guide: Ubuntu

    “# apt dist-upgrade”

    — NVIDIA Driver Installation Guide: Ubuntu

  5. DOCA-Host is NVIDIA's host software for ConnectX and BlueField: drivers, libraries and tools. OFED is the older InfiniBand and RDMA driver stack it replaced. Upgrading from 2.5.x needs a full uninstall first. From 2.6.0 or later, apt can upgrade in place.

    What NVIDIA says (2)

    “To upgrade from DOCA version 2.5.x, all DOCA and OFED related packages should be removed.”

    — DOCA: DOCA-Host Installation and Upgrade

    “Full uninstallation is required before installing DOCA-Host:”

    — DOCA: DOCA-Host Installation and Upgrade

  6. openibd is the service that loads the InfiniBand and RDMA drivers. MST (Mellanox Software Tools) creates device files used by the firmware tools. The DOCA guide updates firmware, loads the drivers, and initializes MST.

    What NVIDIA says (4)

    “host# sudo apt install -y doca-all”

    — DOCA: DOCA-Host Installation and Upgrade

    “Update the firmware: host# sudo apt install -y mlnx-fw-updater”

    — DOCA: DOCA-Host Installation and Upgrade

    “Load the drivers: host# sudo /etc/init.d/openibd restart”

    — DOCA: DOCA-Host Installation and Upgrade

    “Initialize MST: host# sudo mst restart”

    — DOCA: DOCA-Host Installation and Upgrade

  7. Removing packages with apt keeps the package database correct. Deleting files by hand does not. The guide purges a list of driver packages. The list includes nvidia-fabricmanager and nvidia-persistenced.

    What NVIDIA says (2)

    “Follow the below steps to properly uninstall the NVIDIA driver from your system.”

    — NVIDIA Driver Installation Guide: Removing the Driver

    “# apt remove --autoremove --purge -V”

    — NVIDIA Driver Installation Guide: Removing the Driver

Key terms: DOCA

Try it: CUDA Stack Verification

Practice 3.4 (7 questions) Objective page

3.5 NVIDIA Container Toolkit

Official objective: “Install the NVIDIA container toolkit.”

Installing the toolkit and pointing Docker or containerd at the NVIDIA runtime.

Key points

  1. The NVIDIA Container Toolkit lets containers use NVIDIA GPUs. 'nvidia-ctk runtime configure' updates /etc/docker/daemon.json so Docker can use the NVIDIA Container Runtime. Then you restart the Docker daemon.

    What NVIDIA says (3)

    “sudo nvidia-ctk runtime configure --runtime = docker”

    — NVIDIA Container Toolkit: Installation Guide

    “The nvidia-ctk command modifies the /etc/docker/daemon.json file on the host. The file is updated so that Docker can use the NVIDIA Container Runtime.”

    — NVIDIA Container Toolkit: Installation Guide

    “Restart the Docker daemon: $ sudo systemctl restart docker”

    — NVIDIA Container Toolkit: Installation Guide

  2. The NVIDIA Container Toolkit lets containers use the host's GPUs. It relies on the host driver. NVIDIA lists installing the driver as a prerequisite and recommends the package manager.

    What NVIDIA says (2)

    “Install the NVIDIA GPU driver for your Linux distribution.”

    — NVIDIA Container Toolkit: Installation Guide

    “NVIDIA recommends installing the driver by using the package manager for your distribution.”

    — NVIDIA Container Toolkit: Installation Guide

Key terms: NVIDIA Container Toolkit

Practice 3.5 (2 questions) Objective page

3.6 GPUs with Docker

Official objective: “Demonstrate how to use NVIDIA GPUs with Docker.”

Running a sample GPU container to prove the whole chain works.

Key points

  1. A sample workload proves the whole chain works: driver, toolkit and Docker. NVIDIA runs nvidia-smi inside an ubuntu container with '--gpus all'. If the container prints the GPU table, Docker can use the GPUs.

    What NVIDIA says (1)

    “sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”

    — NVIDIA Container Toolkit: Running a Sample Workload

  2. A sample workload proves the whole chain works: driver, toolkit, runtime and Docker. If nvidia-smi inside the container lists the GPUs, the setup is right.

    What NVIDIA says (2)

    “you can verify your installation by running a sample workload.”

    — NVIDIA Container Toolkit: Running a Sample Workload

    “sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”

    — NVIDIA Container Toolkit: Running a Sample Workload

Key terms: NVIDIA Container Toolkit

Try it: NGC Container Flow

Practice 3.6 (2 questions) Objective page

3.7 NGC CLI on hosts

Official objective: “Install NGC CLI on hosts.”

Installing the NGC command line tool and authenticating it with an API key.

Key points

  1. The NGC CLI is NVIDIA's command-line tool for the NGC catalog of containers and models. NVIDIA's examples authenticate it with an NGC API key: run 'ngc config set' and paste the key at the API_KEY prompt.

    What NVIDIA says (2)

    “Here are some examples of using NGC API keys to authenticate with NGC CLI and Docker CLI”

    — NGC Overview

    “ngc config set Paste your key value at the API_KEY prompt”

    — NGC Overview

  2. The NGC CLI is NVIDIA's command-line tool for the NGC catalog of containers and models. NVIDIA recommends always running the latest version for new features, bug fixes and security updates. The CLI can list its own releases and upgrade itself; the NGC CLI Installers page also has the downloads.

    What NVIDIA says (2)

    “Always use the latest NGC CLI version to access the newest features, bug fixes, performance improvements, and security updates.”

    — NGC Overview

    “or run ngc version list to view the latest releases, then upgrade using the following command: ngc version upgrade”

    — NGC Overview

Key terms: NGC

Practice 3.7 (2 questions) Objective page