Installation and Deployment
31% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization
1.1 The Mission Control toolkit
What NVIDIA Mission Control is, what it builds on, and which services run where.
Key points
Mission Control is NVIDIA's operations software for DGX SuperPOD clusters. It uses Base Command Manager (BCM) for core cluster management. BCM provisions nodes, configures software images and assigns roles.
What NVIDIA says (1)
“NVIDIA Mission Control leverages NVIDIA Base Command Manager (BCM) for foundational cluster-management tasks such as provisioning compute nodes, configuring software images, assigning roles, and general cluster administration.”
Mission Control splits the control plane into admin and user parts. The admin Kubernetes nodes host infrastructure services. AHR finds and recovers failed hardware. AJR restarts interrupted jobs.
What NVIDIA says (1)
“Admin Kubernetes Nodes (x86) - x3: BCM-integrated infrastructure services including Observability Stack, Autonomous Hardware Recovery (AHR), and Autonomous Job Recovery (AJR)”
The head node is the central management server. It stores and deploys OS images, coordinates job scheduling and gathers telemetry from every node.
What NVIDIA says (2)
“Provisioning : Centrally store and deploy OS images of the compute, management nodes, and other various services.”
“Metrics : System monitoring and reporting that gather all telemetry from each of the nodes.”
Key terms: NVIDIA Mission Control Base Command Manager Head node
1.2 Monitoring with Base View
How to reach Base View and read cluster health and GPU use at a glance.
Key points
Base View is the browser interface to BCM. cmsh is the command-line interface to the same system. Both talk to CMDaemon, BCM's management daemon.
What NVIDIA says (2)
“Base View is the web application front end to cluster management in BCM.”
“https://<host name or IP address>:8081/base-view”
The overview page is the default view after login. It summarizes the cluster's state, including resources up or down, health checks and GPU usage.
What NVIDIA says (2)
“GPU usage information (includes temperature, power).”
“By default an overview window is displayed, corresponding to the navigation path Cluster > Overview”
Key terms: Base View
1.3 Workload managers in BCM
Adding Slurm or Kubernetes with BCM, and moving nodes between them.
Key points
WLM means workload manager: the scheduler that queues and places jobs. BCM installs it during setup or later with cm-wlm-setup or the Base View wizard.
What NVIDIA says (1)
“A WLM may however also be added and configured after BCM has been installed, by using cm-wlm-setup, which is part of the cm-setup package, or by using the Base View WLM wizard.”
A category is a BCM group of nodes that share configuration. Moving a node between workload managers is a category change. After the reboot, Slurm drops the node from its partitions automatically.
What NVIDIA says (2)
“Essentially, all we do is switch a compute nodes category and reboot.”
“Add the compute node to the proper Run:ai GPU worker category”
Key terms: Workload manager Node category
1.4 Patches, firmware and image sync
Updating software images safely, pushing them to nodes, and firmware through Redfish.
Key points
imageupdate syncs a running node with its software image. By default it only shows what would change. The -w option writes the changes.
What NVIDIA says (1)
“Performing dry run (use synclog command to review result, then pass -w to perform real update)”
A software image is the directory tree that BCM copies onto nodes. cm-chroot-sw-img enters it as a chroot and mounts the special directories that package scripts need.
What NVIDIA says (2)
“Therefore the BCM utility, cm-chroot-sw-img, is strongly recommended to take care of this.”
“root@basecm11:~# cm-chroot-sw-img /cm/images/default-image”
Redfish is the industry standard API for managing servers through their BMC. BCM's firmware command can update the main BIOS and also subsystems such as NICs.
What NVIDIA says (2)
“the modern way of managing BIOS and firmware is with the Redfish standard.”
“Firmware updated via Redfish need not be just the PC main system BIOS, but can also be the flashable software of subsystems, for example: NICs.”
Key terms: Base Command Manager Software image imageupdate Redfish
1.5 Users, roles and access
Where BCM keeps users, how to add one so they can log in, and where to manage them.
Key points
LDAP is a directory service that stores user accounts centrally. BCM runs one on the head nodes, so a user added once works across the cluster. You can also connect an external LDAP server.
What NVIDIA says (2)
“Out of the box, BCM runs its own LDAP service to help manage users and groups.”
“This centralized LDAP service runs on the head nodes of the BCM managed cluster.”
cmsh holds changes locally until you commit them. Users without a password also cannot log in.
What NVIDIA says (2)
“Whenever any changes are made using cmsh , it is important to remember to commit them or else they will not go into effect.”
“Users with unset passwords cannot log in.”
Base View groups user and group management under Identity Management. cmsh and Base View give the same results. One is a CLI, the other a GUI.
What NVIDIA says (2)
“Within Base View, follow the navigation path Identity-Management > Users to manage users.”
“Using cmsh or Base View to manage users and groups will provide the same results.”
1.6 Node, DPU and switch networking
BMC interfaces, node interfaces and DPU settings in cmsh.
Key points
A BMC (baseboard management controller) is the out-of-band management chip in each server. BCM manages it through an interface object of type bmc on the BMC network.
What NVIDIA says (1)
“Once the network has been created, all nodes must be assigned a BMC interface, of type bmc, on this network.”
cmsh is organized into modes. Network interfaces belong to a device, so you enter device mode, select the node and open interfaces.
What NVIDIA says (1)
“The interfaces submode is accessible from the device mode.”
A DPU (data processing unit) is a BlueField card with its own Arm cores. BCM categories have a dpusettings submode next to BIOS, BMC and GPU settings.
What NVIDIA says (1)
“dpusettings ................... Enter DPU settings setup mode”
Key terms: cmsh Baseboard management controller Data processing unit
1.7 Diagnosing cluster issues
Health checks and the cm-diagnose tool for support cases.
Key points
cm-diagnose collects cluster data that helps diagnose issues. You can narrow its options for targeted collection.
What NVIDIA says (1)
“The diagnostic utility cm-diagnose is run from the head node. It gathers data on the cluster that may help diagnose issues.”
Health checks are BCM tests that report PASS or FAIL for a node or service. latesthealthdata lists the latest results.
What NVIDIA says (1)
“In cmsh, the statuses of the services are listed by running the latesthealthdata command (section 10.6.3) from device mode.”
Key terms: Health check cm-diagnose
1.8 Node categories
Grouping nodes that share configuration and software images.
Key points
Categories let you manage many nodes at once. Nodes are usually grouped by hardware type and by role.
What NVIDIA says (2)
“A node category is a group of regular nodes that share the same configuration.”
“Nodes are typically divided into categories based on hardware specifications and their specific purpose.”
A category points to a software image. Several categories can point to the same one.
What NVIDIA says (1)
“categories can share software images. It is not a one-to-one mapping between categories and software images.”
Cloning makes a full copy of the image under a new name. You can switch the category back if the new image has problems.
What NVIDIA says (1)
“It is always good to clone the software images you have for backups.”
Key terms: Software image Node category
1.9 Usage reports
Classifying job metrics and letting a team lead see team data.
Key points
BCM's workload accounting and reporting drills down into job metrics. You can classify them by user, job, account or job name.
What NVIDIA says (2)
“runs jobs, and a job metric can be classified by: • user • job (job ID) • account • job name”
“The classification can be carried out singly. However, it can also be carried out at the same time, like filters.”
Access to job accounting is controlled by tokens in each user's profile. A project manager can see the job data of the users she manages.
What NVIDIA says (1)
“alice can be made the project manager of bob and charlie. This allows her access to the job data of her subordinates”
Key terms: Project manager
1.10 Kubernetes with BCM
Installing Kubernetes with cm-kubernetes-setup, the GPU Operator and etcd sizing.
Key points
A TUI is a text-based menu interface in the terminal. cm-kubernetes-setup also has a command-line mode for automated installs.
What NVIDIA says (1)
“The usual, and recommended way, to install Kubernetes with BCM is to install it interactively from a TUI session”
The GPU Operator installs the device plugin that advertises GPUs to Kubernetes. Without it, no node offers GPUs, so GPU pods cannot be scheduled.
What NVIDIA says (2)
“The NVIDIA GPU Operator is not selected by default.”
“A cluster with NVIDIA GPUs therefore cannot run GPU workloads until the operator is installed.”
etcd is the key-value store that holds Kubernetes state. It needs a majority to work, so an odd count avoids ties. Three nodes avoid a single point of failure.
What NVIDIA says (2)
“An etcd cluster—the Kubernetes distributed key-value storage—runs on an odd number (1, 3, 5 ...) of nodes.”
“a minimum of three nodes is recommended for etcd”
Key terms: etcd NVIDIA GPU Operator
1.11 DOCA services on the DPU
Deploying service containers on BlueField Arm cores and debugging the kubelet.
Key points
DOCA services are containerized programs that run on the DPU. On the DPU, a standalone kubelet watches /etc/kubelet.d and starts a pod for each YAML file.
What NVIDIA says (2)
“cp doca_firefly.yaml /etc/kubelet.d”
“Kubelet automatically pulls the container image from NGC and spawns a pod that runs the container.”
The kubelet log records why a pod failed to start, such as a bad YAML file or missing huge pages. crictl pods lists the pods that did start.
What NVIDIA says (2)
“journalctl -u kubelet Examines the Kubelet logs. Useful when a pod/container fails to spawn.”
“crictl pods Displays currently active K8S pods, and their IDs”
Key terms: Data processing unit Kubelet DOCA service
1.12 Installing Run:ai
Prerequisites and the BCM tool for installing NVIDIA Run:ai.
Key points
Run:ai is NVIDIA's GPU orchestration platform on Kubernetes. Mission Control installs it through BCM's Kubernetes setup tool.
What NVIDIA says (1)
“The installation of NVIDIA Run:ai is done through the cm-kubernetes-setup tool included with BCM 11.”
TLS certificates secure HTTPS. An FQDN is the cluster's full domain name. The Run:ai cluster needs a trusted certificate for that name.
What NVIDIA says (1)
“TLS Certificate must be trusted. Self-signed certificates are not supported.”
The BCM setup assistant uses categories to know which nodes run Kubernetes system services and which run GPU work.
What NVIDIA says (2)
“Before installing NVIDIA Run:ai, make sure BCM node categories are created for:”
“NVIDIA Run:ai GPU worker nodes (for example, dgx-gb200-k8s)”
Key terms: NVIDIA Run:ai Trusted TLS certificate
1.13 Installing Slurm
cm-wlm-setup, NVLink-aware topology and package updates.
Key points
cm-wlm-setup can run as a guided text menu (TUI) or with command-line options. NVIDIA recommends the TUI.
What NVIDIA says (1)
“The recommended way to run the cm-wlm-setup utility is without options or arguments, in which case a TUI dialog starts up.”
A topology plugin tells Slurm how nodes are connected, so it can place a job's nodes close together. BCM 11 adds topology/block and writes topology.conf for you.
What NVIDIA says (1)
“BCM 11 introduces support for topology/block. This is required for BCM to enable NVLINK-aware scheduling of Slurm jobs for GB200/GB300 systems.”
BCM ships Slurm as packages. Minor updates use the normal package manager. Bigger version changes may need manual configuration adjustments.
What NVIDIA says (1)
“WLMs that are packaged with BCM (Slurm) can have their packages updated using standard package update commands (yum update and similar).”
Key terms: Workload manager Slurm topology/block