2.1 Administering Slurm
Draining nodes, partitions, access lists, health checks and IMEX.
Key points
Draining lets running jobs finish but stops new ones. Resume puts the node back into service.
What NVIDIA says (2)
“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”
“Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”
Partitions let you give different jobs different rules, such as run time limits, priority and who may use them.
What NVIDIA says (1)
“A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”
In BCM, Slurm partitions are jobqueue objects. append adds to a list. set replaces the whole list.
What NVIDIA says (2)
“append allowaccounts <user_id>”
“Adding a user or group to a jobqueue (partition):”
A prejob health check runs just before a job starts. If it fails, the workload manager keeps the job off that node.
What NVIDIA says (1)
“A node that has failed a prejob health check is not allowed to run a job.”
IMEX (Internode Memory Exchange) maps GPU memory over NVLink between nodes in an NVLink domain. Per-job IMEX starts just before the job, only on its nodes, which keeps jobs isolated.
What NVIDIA says (1)
“Running the IMEX daemon per job has the advantage that one user running a job cannot read the memory of a job run by another user on another node.”
Key terms
- Health check: A BCM test that marks a node fit or unfit; a node failing a prejob check is kept from running the job.
- Slurm: An open-source workload manager that queues batch jobs and allocates nodes and GPUs to them.
- Slurm partition: A Slurm job queue that groups nodes and sets limits for them.
- Drain: A Slurm node state that lets running jobs finish but accepts no new ones.
- IMEX daemon: The service that lets GPUs in an NVLink domain share memory across nodes; on GB200 it can run per job.
Try it
Sample question
A Slurm node needs maintenance. How do you stop new jobs from landing on it, then return it later?
Show the answer
Answer: scontrol update nodename=<node> state=drain reason="maintenance", then state=resume
Draining lets running jobs finish but stops new ones. Resume puts the node back into service.
What NVIDIA says (2)
“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”
“Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”