3.6 Sharing resources between teams

NCP-AIO · Workload Management (23% of the exam) · Official objective: “Allocate resources between teams with Run:ai, Slurm and Kubernetes”

Quota, over quota, fair share and GPU fractions.

Key points

  1. Over quota means using more than your guaranteed share when GPUs are free. Fairness means that share is returned when its owner needs it.

    What NVIDIA says (1)

    “To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”

    — NVIDIA Run:ai: Over Quota, Fairness and Preemption

  2. Fair share balances scheduling priority by how much each account has used.

    What NVIDIA says (1)

    “Similar to the Slurm command sshare, an administrator can display the Slurm account hierarchy with the fairshare command in cmsh”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  3. Fractions let many small workloads share one GPU, which raises utilization and lets quota be set more precisely.

    What NVIDIA says (1)

    “With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users”

    — NVIDIA Run:ai: GPU Fractions

Key terms

Sample question

In Run:ai, team-a is over its quota and the cluster is full. Team-b, which is under quota, submits a workload. What happens?

Show the answer

Answer: The scheduler preempts some of team-a's over-quota work so team-b gets its share

Over quota means using more than your guaranteed share when GPUs are free. Fairness means that share is returned when its owner needs it.

What NVIDIA says (1)

“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”

— NVIDIA Run:ai: Over Quota, Fairness and Preemption

Practice 3.6 (3 questions) Full Workload Management guide

← 3.5 System management tools · 3.7 Containers from NGC →