1.3 Workload managers in BCM

NCP-AIO · Installation and Deployment (31% of the exam) · Official objective: “Manage job scheduling and resource allocation using BCM’s workload manager (e.g., SLURM or Kubernetes)”

Adding Slurm or Kubernetes with BCM, and moving nodes between them.

Key points

  1. WLM means workload manager: the scheduler that queues and places jobs. BCM installs it during setup or later with cm-wlm-setup or the Base View wizard.

    What NVIDIA says (1)

    “A WLM may however also be added and configured after BCM has been installed, by using cm-wlm-setup, which is part of the cm-setup package, or by using the Base View WLM wizard.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. A category is a BCM group of nodes that share configuration. Moving a node between workload managers is a category change. After the reboot, Slurm drops the node from its partitions automatically.

    What NVIDIA says (2)

    “Essentially, all we do is switch a compute nodes category and reboot.”

    — NVIDIA Mission Control Administration Guide: Adding and Removing Nodes from Run:ai or Slurm

    “Add the compute node to the proper Run:ai GPU worker category”

    — NVIDIA Mission Control Administration Guide: Adding and Removing Nodes from Run:ai or Slurm

Key terms

Sample question

How do you add a WLM such as Slurm to a cluster after BCM is already installed?

Show the answer

Answer: Run cm-wlm-setup, or use the Base View WLM wizard

WLM means workload manager: the scheduler that queues and places jobs. BCM installs it during setup or later with cm-wlm-setup or the Base View wizard.

What NVIDIA says (1)

“A WLM may however also be added and configured after BCM has been installed, by using cm-wlm-setup, which is part of the cm-setup package, or by using the Base View WLM wizard.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Practice 1.3 (2 questions) Full Installation and Deployment guide

← 1.2 Monitoring with Base View · 1.4 Patches, firmware and image sync →