2.1 Administering Slurm

NCP-AIO · Administration (23% of the exam) · Official objective: “Administer Slurm cluster.”

Draining nodes, partitions, access lists, health checks and IMEX.

Key points

  1. Draining lets running jobs finish but stops new ones. Resume puts the node back into service.

    What NVIDIA says (2)

    “scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

    “Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  2. Partitions let you give different jobs different rules, such as run time limits, priority and who may use them.

    What NVIDIA says (1)

    “A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  3. In BCM, Slurm partitions are jobqueue objects. append adds to a list. set replaces the whole list.

    What NVIDIA says (2)

    “append allowaccounts <user_id>”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

    “Adding a user or group to a jobqueue (partition):”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  4. A prejob health check runs just before a job starts. If it fails, the workload manager keeps the job off that node.

    What NVIDIA says (1)

    “A node that has failed a prejob health check is not allowed to run a job.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  5. IMEX (Internode Memory Exchange) maps GPU memory over NVLink between nodes in an NVLink domain. Per-job IMEX starts just before the job, only on its nodes, which keeps jobs isolated.

    What NVIDIA says (1)

    “Running the IMEX daemon per job has the advantage that one user running a job cannot read the memory of a job run by another user on another node.”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

Key terms

Try it

Sample question

A Slurm node needs maintenance. How do you stop new jobs from landing on it, then return it later?

Show the answer

Answer: scontrol update nodename=<node> state=drain reason="maintenance", then state=resume

Draining lets running jobs finish but stops new ones. Resume puts the node back into service.

What NVIDIA says (2)

“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

“Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

Practice 2.1 (5 questions) Full Administration guide

← 1.13 Installing Slurm · 2.2 AI data center architecture →