3.3 Training with Slurm

NCP-AIO · Workload Management (23% of the exam) · Official objective: “Deploy training workloads with Slurm”

Job scripts, GPU requests and containers through Pyxis and Enroot.

Key points

  1. #SBATCH lines in a job script are Slurm options. --gpus sets how many GPUs the job gets.

    What NVIDIA says (1)

    “#SBATCH -p defq #assuming node in defq has a GPU #SBATCH --gpus=1”

    — NVIDIA Base Command Manager 11 User Manual (PDF)

  2. A job script, also called a batch file, holds the Slurm options and the commands to run.

    What NVIDIA says (1)

    “it is usually more convenient to send jobs to Slurm with the sbatch command acting on a job script.”

    — NVIDIA Base Command Manager 11 User Manual (PDF)

  3. Pyxis is a Slurm SPANK plugin. Enroot is the tool that turns container images into unprivileged sandboxes.

    What NVIDIA says (1)

    “The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  4. The job pulls a small image and prints its OS name. If it prints, Pyxis and Enroot are working.

    What NVIDIA says (1)

    “srun --container-image=ubuntu grep PRETTY /etc/os-release pyxis: importing docker image: ubuntu”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms

Try it

Sample question

In a Slurm job script, which directive asks for one GPU?

Show the answer

Answer: #SBATCH --gpus=1

#SBATCH lines in a job script are Slurm options. --gpus sets how many GPUs the job gets.

What NVIDIA says (1)

“#SBATCH -p defq #assuming node in defq has a GPU #SBATCH --gpus=1”

— NVIDIA Base Command Manager 11 User Manual (PDF)

Practice 3.3 (4 questions) Full Workload Management guide

← 3.2 Inference with Run:ai · 3.4 Training with Run:ai →