3.1 BCM install and HA

NCP-AII · Control Plane Installation and Configuration (19% of the exam) · Official objective: “Install Base Command Manager (BCM), configure and verify HA.”

Installing Base Command Manager, licensing it and setting up a failover head node.

Key points

  1. BCM (Base Command Manager) is NVIDIA's cluster manager. It provisions nodes and runs the cluster services. A head node is the server that runs BCM. HA means a second head node can take over if the first one fails. The guide starts HA with the cmha-setup wizard, run as root on the primary head node. The cluster nodes must be powered off first.

    What NVIDIA says (2)

    “Start the cmha-setup CLI wizard as the root user on the primary head node.”

    — DGX SuperPOD Deployment Guide: High Availability

    “The cluster nodes must be powered off before configuring HA.”

    — DGX SuperPOD Deployment Guide: High Availability

  2. In BCM HA, the secondary head node starts as a copy of the primary. You PXE boot it, choose RESCUE, and run /cm/cm-clone-install --failover. PXE (Preboot Execution Environment) means booting from the network instead of a local disk. Later, the Finalize step copies the MySQL database, which holds the cluster configuration.

    What NVIDIA says (2)

    “After the secondary head node has booted into the rescue environment, run the /cm/cm-clone-install --failover command, then enter YES when prompted.”

    — DGX SuperPOD Deployment Guide: High Availability

    “This will clone the MySQL database from the primary to the secondary head node.”

    — DGX SuperPOD Deployment Guide: High Availability

  3. A virtual IP (VIP) is an address that always points at whichever head node is active. Pinging it shows only that one head node answers. cmha status checks failover ping, MySQL and status from both sides. The active head node has an asterisk.

    What NVIDIA says (2)

    “The command tests the configuration from both directions: from the primary head node to the secondary, and from the secondary to the primary. The active head node is indicated by an asterisk.”

    — DGX SuperPOD Deployment Guide: High Availability

    “This will be the IP that should always be used for accessing the active head nodes.”

    — DGX SuperPOD Deployment Guide: High Availability

  4. NAS (network-attached storage) is a file server on the network. Both head nodes must see the same /cm/shared and /home, or a failover would lose files. So cmha-setup copies these directories to the NAS and mounts them everywhere.

    What NVIDIA says (2)

    “cmha-setup will copy the /cm/shared and /home directories to the shared storage and configure both head nodes and all cluster nodes to mount it.”

    — DGX SuperPOD Deployment Guide: High Availability

    “must be stored on an NFS filesystem for HA availability.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

  5. An ISO is a disk image file. The BCM installer ISO can be written to USB or mounted as virtual media through the BMC. After the install finishes and the head node reboots, the guide licenses the cluster with request-license.

    What NVIDIA says (2)

    “License the cluster by running the request-license and providing the product key.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “Ensure that the BIOS of the target head node is configured in UEFI mode and that its boot order is configured to boot the media containing the BCM installer image.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

Key terms

Sample question

You are adding high availability (HA) to a Base Command Manager cluster. Which tool does the DGX SuperPOD deployment guide run on the primary head node to start it?

Show the answer

Answer: cmha-setup, run as root

BCM (Base Command Manager) is NVIDIA's cluster manager. It provisions nodes and runs the cluster services. A head node is the server that runs BCM. HA means a second head node can take over if the first one fails. The guide starts HA with the cmha-setup wizard, run as root on the primary head node. The cluster nodes must be powered off first.

What NVIDIA says (2)

“Start the cmha-setup CLI wizard as the root user on the primary head node.”

— DGX SuperPOD Deployment Guide: High Availability

“The cluster nodes must be powered off before configuring HA.”

— DGX SuperPOD Deployment Guide: High Availability

Practice 3.1 (5 questions) Full Control Plane Installation and Configuration guide

← 2.2 Configuring MIG · 3.2 Installing the OS on nodes →