3.6 Sharing resources between teams
Quota, over quota, fair share and GPU fractions.
Key points
Over quota means using more than your guaranteed share when GPUs are free. Fairness means that share is returned when its owner needs it.
What NVIDIA says (1)
“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”
Fair share balances scheduling priority by how much each account has used.
What NVIDIA says (1)
“Similar to the Slurm command sshare, an administrator can display the Slurm account hierarchy with the fairshare command in cmsh”
Fractions let many small workloads share one GPU, which raises utilization and lets quota be set more precisely.
What NVIDIA says (1)
“With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users”
Key terms
- Deserved quota: The GPU share a Run:ai project is guaranteed; work beyond it is over quota and can be preempted.
- Preemption: The scheduler pausing a lower-priority workload to free its GPUs for another.
- Fair share: Scheduling that balances priority by past use across accounts or projects.
- GPU fractions: A Run:ai feature that splits one GPU memory and compute among several workloads.
Sample question
In Run:ai, team-a is over its quota and the cluster is full. Team-b, which is under quota, submits a workload. What happens?
Show the answer
Answer: The scheduler preempts some of team-a's over-quota work so team-b gets its share
Over quota means using more than your guaranteed share when GPUs are free. Fairness means that share is returned when its owner needs it.
What NVIDIA says (1)
“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”