3.4 Training with Run:ai

NCP-AIO · Workload Management (23% of the exam) · Official objective: “Deploy training workloads with Run:ai”

Priority, preemption, distributed training and checkpoints.

Key points

  1. Preemptible means the scheduler may pause it to free GPUs for more important work. This lets training use spare GPUs beyond the project's quota.

    What NVIDIA says (1)

    “By default, training workloads are assigned a Low priority and are preemptible .”

    — NVIDIA Run:ai: Train Models Using a Standard Training Workload

  2. Distributed training spans nodes, so the workers must coordinate over the network.

    What NVIDIA says (1)

    “Multi-GPU training uses multiple GPUs within a single node, whereas distributed training spans multiple nodes and typically requires coordination between them.”

    — NVIDIA Run:ai: Distributed Training Workloads

  3. A checkpoint is a saved copy of training progress. A resumed workload may land on another node, so local disk may not be there.

    What NVIDIA says (1)

    “Always use shared network storage (e.g., NFS). When a preempted workload is resumed, it may be scheduled on a different node than before.”

    — NVIDIA Run:ai: Checkpointing Preemptible Training Workloads

  4. You choose Workers & master or Workers only. The Master inherits the Worker setup unless you override it.

    What NVIDIA says (1)

    “By default, the Master uses the Worker configuration for shared fields.”

    — NVIDIA Run:ai: Distributed Training Workloads

Key terms

Sample question

By default, what priority and preemption setting does a Run:ai training workload get?

Show the answer

Answer: Low priority and preemptible

Preemptible means the scheduler may pause it to free GPUs for more important work. This lets training use spare GPUs beyond the project's quota.

What NVIDIA says (1)

“By default, training workloads are assigned a Low priority and are preemptible .”

— NVIDIA Run:ai: Train Models Using a Standard Training Workload

Practice 3.4 (4 questions) Full Workload Management guide

← 3.3 Training with Slurm · 3.5 System management tools →