3.2 Orchestration and job scheduling
How Kubernetes, Slurm, Run:ai and Base Command Manager hand out GPUs.
Key points
A device plugin is the Kubernetes mechanism that advertises special hardware to the scheduler. With NVIDIA's device plugin deployed, a container requests GPUs with the nvidia.com/gpu resource type. The scheduler then places the pod only on a node with that many free GPUs.
What NVIDIA says (1)
“With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”
Orchestration is the automated placement and management of workloads on shared resources. NVIDIA Run:ai is a GPU orchestration and optimization platform that aims to maximize compute utilization for AI workloads. Administrators define projects, departments and user roles, and Run:ai enforces policies for fair resource distribution.
What NVIDIA says (2)
“Allows administrators to define Projects, Departments and user roles, enforcing policies for fair resource distribution.”
“NVIDIA Run:ai is a GPU orchestration and optimization platform that helps organizations maximize compute utilization for AI workloads.”
Gang scheduling, also called batch scheduling, treats a group of pods as one unit. The NVIDIA Run:ai Scheduler documentation describes it as a workload of many pods being either fully scheduled or fully pending. That stops half of a job from starting and sitting idle.
What NVIDIA says (1)
“Gang scheduling describes a scheduling principle in which a workload composed of multiple pods is either fully scheduled (all pods are scheduled and running) or fully pending (none of the pods are running).”
Base Command Manager (BCM) is NVIDIA's cluster management software. The DGX SuperPOD reference architecture says it automates provisioning and administration. NVIDIA's product page says it integrates with tools like Slurm or NVIDIA Run:ai. Slurm is a widely used job scheduler for high-performance computing.
What NVIDIA says (3)
“It automates provisioning and administration, and supports clusters into the thousands of nodes”
“NVIDIA Base Command Manager streamlines cluster provisioning, workload management, and infrastructure monitoring.”
“Integrate with tools, like Slurm or NVIDIA Run:ai , to support traditional HPC or AI and analytics workloads across bare-metal and containerized environments.”
Kubernetes is a container orchestration system. An Operator is a Kubernetes pattern for automating the management of a software component. NVIDIA says configuring GPU nodes by hand needs several components, such as drivers and container runtimes, and is error-prone. The GPU Operator automates all of them: the driver, the device plugin, the Container Toolkit, GPU Feature Discovery (GFD) labels and DCGM (Data Center GPU Manager) monitoring.
What NVIDIA says (2)
“The NVIDIA GPU Operator uses the operator framework within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU.”
“These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the NVIDIA Container Toolkit , automatic node labeling using GFD , DCGM based monitoring and others.”
Key terms
- GPU Operator: A Kubernetes operator that installs and manages the NVIDIA driver, device plugin, container toolkit and monitoring.
- Container: A package of an application together with its software dependencies (libraries, configuration, tools) that runs the same on any host.
- Kubernetes: An open-source system that orchestrates containers: it deploys, scales and manages them across a cluster.
- Slurm: An open-source cluster scheduler that queues jobs and assigns them nodes and GPUs.
Try it
Sample question
How does a Kubernetes pod ask for a GPU once the NVIDIA device plugin is running?
Show the answer
Answer: By requesting the nvidia.com/gpu resource in its resource limits, for example nvidia.com/gpu: 1.
A device plugin is the Kubernetes mechanism that advertises special hardware to the scheduler. With NVIDIA's device plugin deployed, a container requests GPUs with the nvidia.com/gpu resource type. The scheduler then places the pod only on a node with that many free GPUs.
What NVIDIA says (1)
“With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”
Practice 3.2 (5 questions) Full AI Operations guide
← 3.1 Data center management and monitoring · 3.3 Monitoring GPUs →