3.3 Cluster setup: categories, Slurm, Enroot, Pyxis

NCP-AII · Control Plane Installation and Configuration (19% of the exam) · Official objective: “Install Cluster (configure category, configure interfaces, install Slurm/Enroot/Pyxis).”

Grouping nodes into categories and installing Slurm with GPU autodetect and container support.

Key points

  1. cmsh is the BCM command-line shell. A category is a group of nodes that share settings, such as the software image and roles. In the guide, DGX H100 nodes are in category dgx-h100. You list nodes and categories with device list.

    What NVIDIA says (3)

    “Check the nodes and their categories.”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “device list -f hostname:20,category:10”

    — DGX SuperPOD Deployment Guide: Initial Cluster Setup

    “the administrator creates a new category called misc. The default category default already exists in a newly installed cluster.”

    — DGX SuperPOD Administration Guide: Provisioning Nodes

  2. Slurm is the workload manager. It queues jobs and assigns nodes and GPUs to them. On a SuperPOD, BCM installs it with the bcm-install-slurm script. Use -A for air-gapped mode, which means a site with no internet access.

    What NVIDIA says (2)

    “Run the bcm-install-slurm script.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

    “Use the -A parameter to run the script in air-gapped mode.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

  3. GRES (generic resources) is how Slurm tracks GPUs. gres.conf lists them. NVML (NVIDIA Management Library) is the driver library that reports GPU details. For DGX H100, the guide sets gpuautodetect to nvml. BCM then writes gres.conf for you.

    What NVIDIA says (3)

    “For DGX H100 systems, generic resources are set to autodetect.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

    “set gpuautodetect nvml”

    — DGX SuperPOD Deployment Guide: Slurm Setup

    “The gres.conf file will be updated automatically by BCM”

    — DGX SuperPOD Deployment Guide: Slurm Setup

  4. SPANK is Slurm's plugin interface. Pyxis is NVIDIA's SPANK plugin. It lets unprivileged users run containers through srun, for example srun --container-image=almalinux:9.

    What NVIDIA says (2)

    “Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”

    — NVIDIA/pyxis: Container plugin for Slurm

    “srun --container-image=almalinux:9 grep PRETTY /etc/os-release”

    — NVIDIA/pyxis: Container plugin for Slurm

  5. A daemon is a background service. Docker uses one. Enroot does not. Enroot imports an image, creates a squashfs file from it, and starts it as an ordinary user. A squashfs is a compressed, read-only file system image.

    What NVIDIA says (3)

    “A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.”

    — NVIDIA/enroot

    “$ enroot import docker://ubuntu $ enroot create ubuntu.sqsh $ enroot start ubuntu”

    — NVIDIA/enroot

    “Standalone (no daemon)”

    — NVIDIA/enroot

Key terms

Try it

Sample question

After BCM setup you want to check which category each node is in. What is a category, and which cmsh command lists it?

Show the answer

Answer: A category is a group of nodes that share configuration. Use device list -f hostname:20,category:10.

cmsh is the BCM command-line shell. A category is a group of nodes that share settings, such as the software image and roles. In the guide, DGX H100 nodes are in category dgx-h100. You list nodes and categories with device list.

What NVIDIA says (3)

“Check the nodes and their categories.”

— DGX SuperPOD Deployment Guide: Initial Cluster Setup

“device list -f hostname:20,category:10”

— DGX SuperPOD Deployment Guide: Initial Cluster Setup

“the administrator creates a new category called misc. The default category default already exists in a newly installed cluster.”

— DGX SuperPOD Administration Guide: Provisioning Nodes

Practice 3.3 (5 questions) Full Control Plane Installation and Configuration guide

← 3.2 Installing the OS on nodes · 3.4 GPU and DOCA drivers →