3.2 Orchestration and job scheduling

NCA-AIIO · AI Operations (22% of the exam) · Official objective: “Describe AI cluster orchestration and job scheduling essentials.”

How Kubernetes, Slurm, Run:ai and Base Command Manager hand out GPUs.

Key points

  1. A device plugin is the Kubernetes mechanism that advertises special hardware to the scheduler. With NVIDIA's device plugin deployed, a container requests GPUs with the nvidia.com/gpu resource type. The scheduler then places the pod only on a node with that many free GPUs.

    What NVIDIA says (1)

    “With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”

    — NVIDIA device plugin for Kubernetes

  2. Orchestration is the automated placement and management of workloads on shared resources. NVIDIA Run:ai is a GPU orchestration and optimization platform that aims to maximize compute utilization for AI workloads. Administrators define projects, departments and user roles, and Run:ai enforces policies for fair resource distribution.

    What NVIDIA says (2)

    “Allows administrators to define Projects, Departments and user roles, enforcing policies for fair resource distribution.”

    — NVIDIA Run:ai Documentation: Overview

    “NVIDIA Run:ai is a GPU orchestration and optimization platform that helps organizations maximize compute utilization for AI workloads.”

    — NVIDIA Run:ai Documentation: Overview

  3. Gang scheduling, also called batch scheduling, treats a group of pods as one unit. The NVIDIA Run:ai Scheduler documentation describes it as a workload of many pods being either fully scheduled or fully pending. That stops half of a job from starting and sitting idle.

    What NVIDIA says (1)

    “Gang scheduling describes a scheduling principle in which a workload composed of multiple pods is either fully scheduled (all pods are scheduled and running) or fully pending (none of the pods are running).”

    — NVIDIA Run:ai Scheduler: Concepts and Principles

  4. Base Command Manager (BCM) is NVIDIA's cluster management software. The DGX SuperPOD reference architecture says it automates provisioning and administration. NVIDIA's product page says it integrates with tools like Slurm or NVIDIA Run:ai. Slurm is a widely used job scheduler for high-performance computing.

    What NVIDIA says (3)

    “It automates provisioning and administration, and supports clusters into the thousands of nodes”

    — DGX SuperPOD H100 Reference Architecture: Components

    “NVIDIA Base Command Manager streamlines cluster provisioning, workload management, and infrastructure monitoring.”

    — NVIDIA Base Command Manager Documentation

    “Integrate with tools, like Slurm or NVIDIA Run:ai , to support traditional HPC or AI and analytics workloads across bare-metal and containerized environments.”

    — NVIDIA Base Command Manager

  5. Kubernetes is a container orchestration system. An Operator is a Kubernetes pattern for automating the management of a software component. NVIDIA says configuring GPU nodes by hand needs several components, such as drivers and container runtimes, and is error-prone. The GPU Operator automates all of them: the driver, the device plugin, the Container Toolkit, GPU Feature Discovery (GFD) labels and DCGM (Data Center GPU Manager) monitoring.

    What NVIDIA says (2)

    “The NVIDIA GPU Operator uses the operator framework within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU.”

    — About the NVIDIA GPU Operator

    “These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the NVIDIA Container Toolkit , automatic node labeling using GFD , DCGM based monitoring and others.”

    — About the NVIDIA GPU Operator

Key terms

Try it

Sample question

How does a Kubernetes pod ask for a GPU once the NVIDIA device plugin is running?

Show the answer

Answer: By requesting the nvidia.com/gpu resource in its resource limits, for example nvidia.com/gpu: 1.

A device plugin is the Kubernetes mechanism that advertises special hardware to the scheduler. With NVIDIA's device plugin deployed, a container requests GPUs with the nvidia.com/gpu resource type. The scheduler then places the pod only on a node with that many free GPUs.

What NVIDIA says (1)

“With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”

— NVIDIA device plugin for Kubernetes

Practice 3.2 (5 questions) Full AI Operations guide

← 3.1 Data center management and monitoring · 3.3 Monitoring GPUs →