Installation and Deployment

31% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization

1.1 The Mission Control toolkit

Official objective: “Describe the Mission Control toolkit”

What NVIDIA Mission Control is, what it builds on, and which services run where.

Key points

  1. Mission Control is NVIDIA's operations software for DGX SuperPOD clusters. It uses Base Command Manager (BCM) for core cluster management. BCM provisions nodes, configures software images and assigns roles.

    What NVIDIA says (1)

    “NVIDIA Mission Control leverages NVIDIA Base Command Manager (BCM) for foundational cluster-management tasks such as provisioning compute nodes, configuring software images, assigning roles, and general cluster administration.”

    — NVIDIA Mission Control Administration Guide: Software Stack

  2. Mission Control splits the control plane into admin and user parts. The admin Kubernetes nodes host infrastructure services. AHR finds and recovers failed hardware. AJR restarts interrupted jobs.

    What NVIDIA says (1)

    “Admin Kubernetes Nodes (x86) - x3: BCM-integrated infrastructure services including Observability Stack, Autonomous Hardware Recovery (AHR), and Autonomous Job Recovery (AJR)”

    — NVIDIA Mission Control Administration Guide: Overview

  3. The head node is the central management server. It stores and deploys OS images, coordinates job scheduling and gathers telemetry from every node.

    What NVIDIA says (2)

    “Provisioning : Centrally store and deploy OS images of the compute, management nodes, and other various services.”

    — NVIDIA Mission Control Administration Guide: Overview

    “Metrics : System monitoring and reporting that gather all telemetry from each of the nodes.”

    — NVIDIA Mission Control Administration Guide: Overview

Key terms: NVIDIA Mission Control Base Command Manager Head node

Practice 1.1 (3 questions)

1.2 Monitoring with Base View

Official objective: “Use BCM’s Base View interface to monitor cluster performance, resource utilization, and node health in real time.”

How to reach Base View and read cluster health and GPU use at a glance.

Key points

  1. Base View is the browser interface to BCM. cmsh is the command-line interface to the same system. Both talk to CMDaemon, BCM's management daemon.

    What NVIDIA says (2)

    “Base View is the web application front end to cluster management in BCM.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

    “https://<host name or IP address>:8081/base-view”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. The overview page is the default view after login. It summarizes the cluster's state, including resources up or down, health checks and GPU usage.

    What NVIDIA says (2)

    “GPU usage information (includes temperature, power).”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

    “By default an overview window is displayed, corresponding to the navigation path Cluster > Overview”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Base View

Practice 1.2 (2 questions)

1.3 Workload managers in BCM

Official objective: “Manage job scheduling and resource allocation using BCM’s workload manager (e.g., SLURM or Kubernetes)”

Adding Slurm or Kubernetes with BCM, and moving nodes between them.

Key points

  1. WLM means workload manager: the scheduler that queues and places jobs. BCM installs it during setup or later with cm-wlm-setup or the Base View wizard.

    What NVIDIA says (1)

    “A WLM may however also be added and configured after BCM has been installed, by using cm-wlm-setup, which is part of the cm-setup package, or by using the Base View WLM wizard.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. A category is a BCM group of nodes that share configuration. Moving a node between workload managers is a category change. After the reboot, Slurm drops the node from its partitions automatically.

    What NVIDIA says (2)

    “Essentially, all we do is switch a compute nodes category and reboot.”

    — NVIDIA Mission Control Administration Guide: Adding and Removing Nodes from Run:ai or Slurm

    “Add the compute node to the proper Run:ai GPU worker category”

    — NVIDIA Mission Control Administration Guide: Adding and Removing Nodes from Run:ai or Slurm

Key terms: Workload manager Node category

Practice 1.3 (2 questions)

1.4 Patches, firmware and image sync

Official objective: “Apply patches, update firmware, and synchronize software images across cluster nodes using BCM”

Updating software images safely, pushing them to nodes, and firmware through Redfish.

Key points

  1. imageupdate syncs a running node with its software image. By default it only shows what would change. The -w option writes the changes.

    What NVIDIA says (1)

    “Performing dry run (use synclog command to review result, then pass -w to perform real update)”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. A software image is the directory tree that BCM copies onto nodes. cm-chroot-sw-img enters it as a chroot and mounts the special directories that package scripts need.

    What NVIDIA says (2)

    “Therefore the BCM utility, cm-chroot-sw-img, is strongly recommended to take care of this.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

    “root@basecm11:~# cm-chroot-sw-img /cm/images/default-image”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  3. Redfish is the industry standard API for managing servers through their BMC. BCM's firmware command can update the main BIOS and also subsystems such as NICs.

    What NVIDIA says (2)

    “the modern way of managing BIOS and firmware is with the Redfish standard.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

    “Firmware updated via Redfish need not be just the PC main system BIOS, but can also be the flashable software of subsystems, for example: NICs.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Base Command Manager Software image imageupdate Redfish

Practice 1.4 (3 questions)

1.5 Users, roles and access

Official objective: “Administer user accounts, roles, and permissions to ensure secure access to the cluster using BCM”

Where BCM keeps users, how to add one so they can log in, and where to manage them.

Key points

  1. LDAP is a directory service that stores user accounts centrally. BCM runs one on the head nodes, so a user added once works across the cluster. You can also connect an external LDAP server.

    What NVIDIA says (2)

    “Out of the box, BCM runs its own LDAP service to help manage users and groups.”

    — NVIDIA Mission Control Administration Guide: Software Stack

    “This centralized LDAP service runs on the head nodes of the BCM managed cluster.”

    — NVIDIA Mission Control Administration Guide: Software Stack

  2. cmsh holds changes locally until you commit them. Users without a password also cannot log in.

    What NVIDIA says (2)

    “Whenever any changes are made using cmsh , it is important to remember to commit them or else they will not go into effect.”

    — NVIDIA Mission Control Administration Guide: Software Stack

    “Users with unset passwords cannot log in.”

    — NVIDIA Mission Control Administration Guide: Software Stack

  3. Base View groups user and group management under Identity Management. cmsh and Base View give the same results. One is a CLI, the other a GUI.

    What NVIDIA says (2)

    “Within Base View, follow the navigation path Identity-Management > Users to manage users.”

    — NVIDIA Mission Control Administration Guide: Software Stack

    “Using cmsh or Base View to manage users and groups will provide the same results.”

    — NVIDIA Mission Control Administration Guide: Software Stack

Key terms: Base View cmsh LDAP

Practice 1.5 (3 questions)

1.6 Node, DPU and switch networking

Official objective: “Configure and monitor network settings for cluster nodes, DPUs, and switches using BCM”

BMC interfaces, node interfaces and DPU settings in cmsh.

Key points

  1. A BMC (baseboard management controller) is the out-of-band management chip in each server. BCM manages it through an interface object of type bmc on the BMC network.

    What NVIDIA says (1)

    “Once the network has been created, all nodes must be assigned a BMC interface, of type bmc, on this network.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. cmsh is organized into modes. Network interfaces belong to a device, so you enter device mode, select the node and open interfaces.

    What NVIDIA says (1)

    “The interfaces submode is accessible from the device mode.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  3. A DPU (data processing unit) is a BlueField card with its own Arm cores. BCM categories have a dpusettings submode next to BIOS, BMC and GPU settings.

    What NVIDIA says (1)

    “dpusettings ................... Enter DPU settings setup mode”

    — NVIDIA Mission Control Administration Guide: Node and Category Management

Key terms: cmsh Baseboard management controller Data processing unit

Practice 1.6 (3 questions)

1.7 Diagnosing cluster issues

Official objective: “Diagnose and resolve cluster issues, such as job failures, node outages, or resource bottlenecks, using BCM.”

Health checks and the cm-diagnose tool for support cases.

Key points

  1. cm-diagnose collects cluster data that helps diagnose issues. You can narrow its options for targeted collection.

    What NVIDIA says (1)

    “The diagnostic utility cm-diagnose is run from the head node. It gathers data on the cluster that may help diagnose issues.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. Health checks are BCM tests that report PASS or FAIL for a node or service. latesthealthdata lists the latest results.

    What NVIDIA says (1)

    “In cmsh, the statuses of the services are listed by running the latesthealthdata command (section 10.6.3) from device mode.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Health check cm-diagnose

Practice 1.7 (2 questions)

1.8 Node categories

Official objective: “Use BCM to organize and configure compute nodes into categories based on hardware or workload requirements.”

Grouping nodes that share configuration and software images.

Key points

  1. Categories let you manage many nodes at once. Nodes are usually grouped by hardware type and by role.

    What NVIDIA says (2)

    “A node category is a group of regular nodes that share the same configuration.”

    — NVIDIA Mission Control Administration Guide: Node and Category Management

    “Nodes are typically divided into categories based on hardware specifications and their specific purpose.”

    — NVIDIA Mission Control Administration Guide: Node and Category Management

  2. A category points to a software image. Several categories can point to the same one.

    What NVIDIA says (1)

    “categories can share software images. It is not a one-to-one mapping between categories and software images.”

    — NVIDIA Mission Control Administration Guide: Node and Category Management

  3. Cloning makes a full copy of the image under a new name. You can switch the category back if the new image has problems.

    What NVIDIA says (1)

    “It is always good to clone the software images you have for backups.”

    — NVIDIA Mission Control Administration Guide: Node and Category Management

Key terms: Software image Node category

Practice 1.8 (3 questions)

1.9 Usage reports

Official objective: “Using BCM, maintain documentation and generate reports on cluster usage, performance, and issues.”

Classifying job metrics and letting a team lead see team data.

Key points

  1. BCM's workload accounting and reporting drills down into job metrics. You can classify them by user, job, account or job name.

    What NVIDIA says (2)

    “runs jobs, and a job metric can be classified by: • user • job (job ID) • account • job name”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

    “The classification can be carried out singly. However, it can also be carried out at the same time, like filters.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. Access to job accounting is controlled by tokens in each user's profile. A project manager can see the job data of the users she manages.

    What NVIDIA says (1)

    “alice can be made the project manager of bob and charlie. This allows her access to the job data of her subordinates”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Project manager

Practice 1.9 (2 questions)

1.10 Kubernetes with BCM

Official objective: “Install and initialize Kubernetes on NVIDIA hosts using BCM”

Installing Kubernetes with cm-kubernetes-setup, the GPU Operator and etcd sizing.

Key points

  1. A TUI is a text-based menu interface in the terminal. cm-kubernetes-setup also has a command-line mode for automated installs.

    What NVIDIA says (1)

    “The usual, and recommended way, to install Kubernetes with BCM is to install it interactively from a TUI session”

    — NVIDIA Base Command Manager 11 Containerization Manual (PDF)

  2. The GPU Operator installs the device plugin that advertises GPUs to Kubernetes. Without it, no node offers GPUs, so GPU pods cannot be scheduled.

    What NVIDIA says (2)

    “The NVIDIA GPU Operator is not selected by default.”

    — NVIDIA Base Command Manager 11 Containerization Manual (PDF)

    “A cluster with NVIDIA GPUs therefore cannot run GPU workloads until the operator is installed.”

    — NVIDIA Base Command Manager 11 Containerization Manual (PDF)

  3. etcd is the key-value store that holds Kubernetes state. It needs a majority to work, so an odd count avoids ties. Three nodes avoid a single point of failure.

    What NVIDIA says (2)

    “An etcd cluster—the Kubernetes distributed key-value storage—runs on an odd number (1, 3, 5 ...) of nodes.”

    — NVIDIA Base Command Manager 11 Containerization Manual (PDF)

    “a minimum of three nodes is recommended for etcd”

    — NVIDIA Base Command Manager 11 Containerization Manual (PDF)

Key terms: etcd NVIDIA GPU Operator

Practice 1.10 (3 questions)

1.11 DOCA services on the DPU

Official objective: “Deploy DOCA Services on DPU Arm”

Deploying service containers on BlueField Arm cores and debugging the kubelet.

Key points

  1. DOCA services are containerized programs that run on the DPU. On the DPU, a standalone kubelet watches /etc/kubelet.d and starts a pod for each YAML file.

    What NVIDIA says (2)

    “cp doca_firefly.yaml /etc/kubelet.d”

    — DOCA Container Deployment Guide

    “Kubelet automatically pulls the container image from NGC and spawns a pod that runs the container.”

    — DOCA Container Deployment Guide

  2. The kubelet log records why a pod failed to start, such as a bad YAML file or missing huge pages. crictl pods lists the pods that did start.

    What NVIDIA says (2)

    “journalctl -u kubelet Examines the Kubelet logs. Useful when a pod/container fails to spawn.”

    — BlueField Troubleshooting Guide: DOCA Services

    “crictl pods Displays currently active K8S pods, and their IDs”

    — BlueField Troubleshooting Guide: DOCA Services

Key terms: Data processing unit Kubelet DOCA service

Practice 1.11 (2 questions)

1.12 Installing Run:ai

Official objective: “Install Run:ai”

Prerequisites and the BCM tool for installing NVIDIA Run:ai.

Key points

  1. Run:ai is NVIDIA's GPU orchestration platform on Kubernetes. Mission Control installs it through BCM's Kubernetes setup tool.

    What NVIDIA says (1)

    “The installation of NVIDIA Run:ai is done through the cm-kubernetes-setup tool included with BCM 11.”

    — NVIDIA Mission Control Administration Guide: Run:ai Installation

  2. TLS certificates secure HTTPS. An FQDN is the cluster's full domain name. The Run:ai cluster needs a trusted certificate for that name.

    What NVIDIA says (1)

    “TLS Certificate must be trusted. Self-signed certificates are not supported.”

    — NVIDIA Run:ai: Install Using Base Command Manager

  3. The BCM setup assistant uses categories to know which nodes run Kubernetes system services and which run GPU work.

    What NVIDIA says (2)

    “Before installing NVIDIA Run:ai, make sure BCM node categories are created for:”

    — NVIDIA Mission Control Administration Guide: Run:ai Installation

    “NVIDIA Run:ai GPU worker nodes (for example, dgx-gb200-k8s)”

    — NVIDIA Mission Control Administration Guide: Run:ai Installation

Key terms: NVIDIA Run:ai Trusted TLS certificate

Practice 1.12 (3 questions)

1.13 Installing Slurm

Official objective: “Install Slurm”

cm-wlm-setup, NVLink-aware topology and package updates.

Key points

  1. cm-wlm-setup can run as a guided text menu (TUI) or with command-line options. NVIDIA recommends the TUI.

    What NVIDIA says (1)

    “The recommended way to run the cm-wlm-setup utility is without options or arguments, in which case a TUI dialog starts up.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. A topology plugin tells Slurm how nodes are connected, so it can place a job's nodes close together. BCM 11 adds topology/block and writes topology.conf for you.

    What NVIDIA says (1)

    “BCM 11 introduces support for topology/block. This is required for BCM to enable NVLINK-aware scheduling of Slurm jobs for GB200/GB300 systems.”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  3. BCM ships Slurm as packages. Minor updates use the normal package manager. Bigger version changes may need manual configuration adjustments.

    What NVIDIA says (1)

    “WLMs that are packaged with BCM (Slurm) can have their packages updated using standard package update commands (yum update and similar).”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Workload manager Slurm topology/block

Practice 1.13 (3 questions)