Administration

23% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization

2.1 Administering Slurm

Official objective: “Administer Slurm cluster.”

Draining nodes, partitions, access lists, health checks and IMEX.

Key points

  1. Draining lets running jobs finish but stops new ones. Resume puts the node back into service.

    What NVIDIA says (2)

    “scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

    “Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  2. Partitions let you give different jobs different rules, such as run time limits, priority and who may use them.

    What NVIDIA says (1)

    “A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  3. In BCM, Slurm partitions are jobqueue objects. append adds to a list. set replaces the whole list.

    What NVIDIA says (2)

    “append allowaccounts <user_id>”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

    “Adding a user or group to a jobqueue (partition):”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

  4. A prejob health check runs just before a job starts. If it fails, the workload manager keeps the job off that node.

    What NVIDIA says (1)

    “A node that has failed a prejob health check is not allowed to run a job.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  5. IMEX (Internode Memory Exchange) maps GPU memory over NVLink between nodes in an NVLink domain. Per-job IMEX starts just before the job, only on its nodes, which keeps jobs isolated.

    What NVIDIA says (1)

    “Running the IMEX daemon per job has the advantage that one user running a job cannot read the memory of a job run by another user on another node.”

    — NVIDIA Mission Control Administration Guide: Slurm Workload Management

Key terms: Health check Slurm Slurm partition Drain IMEX daemon

Try it: Slurm Scheduler

Practice 2.1 (5 questions)

2.2 AI data center architecture

Official objective: “Describe data center architecture for AI Workloads”

The SuperPOD fabrics, storage, management network and Multi-Node NVLink.

Key points

  1. A fabric is a network of switches and cables. Storage reads and writes would otherwise compete with GPU-to-GPU traffic.

    What NVIDIA says (1)

    “Storage traffic is dedicated to its own fabric to remove interference with the node-to-node application traffic that can degrade overall performance.”

    — NVIDIA Mission Control Administration Guide: Overview

  2. Out-of-band (OOB) means separate from the data networks. It reaches the management controllers even when an OS is down.

    What NVIDIA says (1)

    “The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”

    — NVIDIA Mission Control Administration Guide: Overview

  3. HSS is shared, high-bandwidth storage for every node. Home directories use a separate NFS share.

    What NVIDIA says (1)

    “High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”

    — NVIDIA Mission Control Administration Guide: Overview

  4. Users do not log in to GPU nodes directly. They work on login nodes, which are Slurm clients with the file systems mounted.

    What NVIDIA says (1)

    “Entry point to the DGX SuperPOD for users. CPU-based nodes that are Slurm clients with filesystems mounted to support development, job submission, job monitoring, and file management.”

    — NVIDIA Mission Control Administration Guide: Overview

  5. NVLink connects GPUs directly at high speed. On GB200 and GB300, NVLink switches extend this across trays in a rack.

    What NVIDIA says (1)

    “Multi-Node NVLink is a capability enabled over an NVLink Switch network where multiple systems are interconnected to form a large GPU memory fabric also known as an NVLink Domain.”

    — NVIDIA Mission Control Administration Guide: Overview

Key terms: Baseboard management controller DGX SuperPOD Out-of-band network High-speed storage Multi-Node NVLink

Practice 2.2 (5 questions)

2.3 Administering Run:ai

Official objective: “Administer Run:ai”

Projects, departments, roles and node pools.

Key points

  1. A project can be a team, a person or an initiative. In Kubernetes, each project becomes a namespace.

    What NVIDIA says (2)

    “NVIDIA Run:ai uses Projects as the primary organization management unit.”

    — NVIDIA Run:ai: Projects

    “Projects are manifested as Kubernetes namespaces.”

    — NVIDIA Run:ai: Projects

  2. Departments sit above projects. They let you allocate quota across projects and apply policies at department level.

    What NVIDIA says (1)

    “Departments group multiple projects under a shared organizational scope.”

    — NVIDIA Run:ai: Departments

  3. A subject is a user, group or service account. A scope is the part of the organization the role applies to. Run:ai has predefined roles and allows custom ones.

    What NVIDIA says (1)

    “A role defines a set of permissions that can be assigned to a subject in a scope”

    — NVIDIA Run:ai: Roles

  4. A node pool groups nodes by a label, such as GPU type. With over quota enabled, projects can still use a new pool before getting quota there.

    What NVIDIA says (1)

    “Once created, the new node pool is automatically assigned to all projects and departments with a quota of zero GPU resources, unlimited CPU resources, and over quota enabled”

    — NVIDIA Run:ai: Node Pools

  5. Each node pool has its own scheduler instance. Workloads sent to a pool are scheduled by that instance.

    What NVIDIA says (1)

    “Creating a new node pool creates a new instance of the NVIDIA Run:ai Scheduler .”

    — NVIDIA Run:ai: Node Pools

Key terms: NVIDIA Run:ai Run:ai project Run:ai department Node pool

Practice 2.3 (5 questions)

2.4 Administering Kubernetes for GPUs

Official objective: “Administer Kubernetes”

Checking the GPU Operator, NFD and GPU time-slicing.

Key points

  1. The GPU Operator runs several pods: driver, toolkit, device plugin, feature discovery and validators. All should be Running or Completed.

    What NVIDIA says (1)

    “Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”

    — GPU Operator: Installing the NVIDIA GPU Operator

  2. NFD (Node Feature Discovery) labels nodes with their hardware features. The Operator needs it on every node. If NFD already runs, deploying a second copy must be turned off.

    What NVIDIA says (1)

    “If NFD is already running in the cluster, then you must disable deploying NFD when you install the Operator.”

    — GPU Operator: Installing the NVIDIA GPU Operator

  3. Time-slicing lets several pods take turns on one GPU. It shares the GPU but does not partition memory. MIG does.

    What NVIDIA says (1)

    “Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”

    — Time-Slicing GPUs in Kubernetes

  4. The device plugin tells Kubernetes how many GPUs a node has. With time-slicing, it advertises several replicas of each GPU.

    What NVIDIA says (1)

    “The NVIDIA GPU Operator enables oversubscription of GPUs through a set of extended options for the NVIDIA Kubernetes Device Plugin .”

    — Time-Slicing GPUs in Kubernetes

Key terms: NVIDIA GPU Operator Node Feature Discovery GPU time-slicing

Try it: Kubernetes GPU Ops

Practice 2.4 (4 questions)

2.5 Configuring MIG

Official objective: “Configure MIG”

MIG strategies, profiles, creating instances and targeting one MIG device.

Key points

  1. MIG Manager is the GPU Operator component that applies MIG layouts. It watches the node label and reconfigures the GPUs.

    What NVIDIA says (2)

    “nvidia.com/mig.config=all-1g.10gb”

    — GPU Operator with MIG

    “GPU Operator deploys MIG Manager to manage MIG configuration on nodes in your Kubernetes cluster.”

    — GPU Operator with MIG

  2. The strategy tells the Operator how MIG devices are exposed. You choose it at install time, before you configure MIG.

    What NVIDIA says (2)

    “single : MIG mode is enabled on all GPUs on a node.”

    — GPU Operator with MIG

    “mixed : MIG mode is not enabled on all GPUs on a node.”

    — GPU Operator with MIG

  3. A GPU instance profile is a fixed slice size, such as 1g.5gb. -lgip lists profiles with free and total counts.

    What NVIDIA says (1)

    “$ nvidia-smi mig -lgip”

    — MIG User Guide: Getting Started with MIG

  4. -cgi creates GPU instances. You can name a profile by ID (9) or by name (3g.20gb). -C also creates the matching compute instances.

    What NVIDIA says (2)

    “sudo nvidia-smi mig -cgi 9,3g.20gb -C”

    — MIG User Guide: Getting Started with MIG

    “Successfully created GPU instance ID 2 on GPU 0 using profile MIG 3g.20gb (ID 9)”

    — MIG User Guide: Getting Started with MIG

  5. Each MIG device has its own UUID. CUDA_VISIBLE_DEVICES limits an app to the devices you name.

    What NVIDIA says (2)

    “$ nvidia-smi -L GPU 0: A100-SXM4-40GB”

    — MIG User Guide: Getting Started with MIG

    “$ CUDA_VISIBLE_DEVICES=MIG-c7384736-a75d-5afc-978f-d2f1294409fd ./BlackScholes &”

    — MIG User Guide: Getting Started with MIG

Key terms: Multi-Instance GPU

Try it: MIG Partitioning

Practice 2.5 (5 questions)