3.4 Training with Run:ai
Priority, preemption, distributed training and checkpoints.
Key points
Preemptible means the scheduler may pause it to free GPUs for more important work. This lets training use spare GPUs beyond the project's quota.
What NVIDIA says (1)
“By default, training workloads are assigned a Low priority and are preemptible .”
Distributed training spans nodes, so the workers must coordinate over the network.
What NVIDIA says (1)
“Multi-GPU training uses multiple GPUs within a single node, whereas distributed training spans multiple nodes and typically requires coordination between them.”
A checkpoint is a saved copy of training progress. A resumed workload may land on another node, so local disk may not be there.
What NVIDIA says (1)
“Always use shared network storage (e.g., NFS). When a preempted workload is resumed, it may be scheduled on a different node than before.”
You choose Workers & master or Workers only. The Master inherits the Worker setup unless you override it.
What NVIDIA says (1)
“By default, the Master uses the Worker configuration for shared fields.”
Key terms
- Preemption: The scheduler pausing a lower-priority workload to free its GPUs for another.
- Checkpoint: A saved copy of training progress so a stopped job can resume.
Sample question
By default, what priority and preemption setting does a Run:ai training workload get?
Show the answer
Answer: Low priority and preemptible
Preemptible means the scheduler may pause it to free GPUs for more important work. This lets training use spare GPUs beyond the project's quota.
What NVIDIA says (1)
“By default, training workloads are assigned a Low priority and are preemptible .”