Administration
23% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization
2.1 Administering Slurm
Draining nodes, partitions, access lists, health checks and IMEX.
Key points
Draining lets running jobs finish but stops new ones. Resume puts the node back into service.
What NVIDIA says (2)
“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”
“Resume nodes similarly: scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = resume”
Partitions let you give different jobs different rules, such as run time limits, priority and who may use them.
What NVIDIA says (1)
“A Slurm partition is a distinct job queue that groups compute nodes together and enables the ability to set specific resource constraints and limits for those nodes.”
In BCM, Slurm partitions are jobqueue objects. append adds to a list. set replaces the whole list.
What NVIDIA says (2)
“append allowaccounts <user_id>”
“Adding a user or group to a jobqueue (partition):”
A prejob health check runs just before a job starts. If it fails, the workload manager keeps the job off that node.
What NVIDIA says (1)
“A node that has failed a prejob health check is not allowed to run a job.”
IMEX (Internode Memory Exchange) maps GPU memory over NVLink between nodes in an NVLink domain. Per-job IMEX starts just before the job, only on its nodes, which keeps jobs isolated.
What NVIDIA says (1)
“Running the IMEX daemon per job has the advantage that one user running a job cannot read the memory of a job run by another user on another node.”
Key terms: Health check Slurm Slurm partition Drain IMEX daemon
Try it: Slurm Scheduler
2.2 AI data center architecture
The SuperPOD fabrics, storage, management network and Multi-Node NVLink.
Key points
A fabric is a network of switches and cables. Storage reads and writes would otherwise compete with GPU-to-GPU traffic.
What NVIDIA says (1)
“Storage traffic is dedicated to its own fabric to remove interference with the node-to-node application traffic that can degrade overall performance.”
Out-of-band (OOB) means separate from the data networks. It reaches the management controllers even when an OS is down.
What NVIDIA says (1)
“The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”
HSS is shared, high-bandwidth storage for every node. Home directories use a separate NFS share.
What NVIDIA says (1)
“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”
Users do not log in to GPU nodes directly. They work on login nodes, which are Slurm clients with the file systems mounted.
What NVIDIA says (1)
“Entry point to the DGX SuperPOD for users. CPU-based nodes that are Slurm clients with filesystems mounted to support development, job submission, job monitoring, and file management.”
NVLink connects GPUs directly at high speed. On GB200 and GB300, NVLink switches extend this across trays in a rack.
What NVIDIA says (1)
“Multi-Node NVLink is a capability enabled over an NVLink Switch network where multiple systems are interconnected to form a large GPU memory fabric also known as an NVLink Domain.”
Key terms: Baseboard management controller DGX SuperPOD Out-of-band network High-speed storage Multi-Node NVLink
2.3 Administering Run:ai
Projects, departments, roles and node pools.
Key points
A project can be a team, a person or an initiative. In Kubernetes, each project becomes a namespace.
What NVIDIA says (2)
“NVIDIA Run:ai uses Projects as the primary organization management unit.”
“Projects are manifested as Kubernetes namespaces.”
Departments sit above projects. They let you allocate quota across projects and apply policies at department level.
What NVIDIA says (1)
“Departments group multiple projects under a shared organizational scope.”
A subject is a user, group or service account. A scope is the part of the organization the role applies to. Run:ai has predefined roles and allows custom ones.
What NVIDIA says (1)
“A role defines a set of permissions that can be assigned to a subject in a scope”
A node pool groups nodes by a label, such as GPU type. With over quota enabled, projects can still use a new pool before getting quota there.
What NVIDIA says (1)
“Once created, the new node pool is automatically assigned to all projects and departments with a quota of zero GPU resources, unlimited CPU resources, and over quota enabled”
Each node pool has its own scheduler instance. Workloads sent to a pool are scheduled by that instance.
What NVIDIA says (1)
“Creating a new node pool creates a new instance of the NVIDIA Run:ai Scheduler .”
Key terms: NVIDIA Run:ai Run:ai project Run:ai department Node pool
2.4 Administering Kubernetes for GPUs
Checking the GPU Operator, NFD and GPU time-slicing.
Key points
The GPU Operator runs several pods: driver, toolkit, device plugin, feature discovery and validators. All should be Running or Completed.
What NVIDIA says (1)
“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”
NFD (Node Feature Discovery) labels nodes with their hardware features. The Operator needs it on every node. If NFD already runs, deploying a second copy must be turned off.
What NVIDIA says (1)
“If NFD is already running in the cluster, then you must disable deploying NFD when you install the Operator.”
Time-slicing lets several pods take turns on one GPU. It shares the GPU but does not partition memory. MIG does.
What NVIDIA says (1)
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”
The device plugin tells Kubernetes how many GPUs a node has. With time-slicing, it advertises several replicas of each GPU.
What NVIDIA says (1)
“The NVIDIA GPU Operator enables oversubscription of GPUs through a set of extended options for the NVIDIA Kubernetes Device Plugin .”
Key terms: NVIDIA GPU Operator Node Feature Discovery GPU time-slicing
Try it: Kubernetes GPU Ops
2.5 Configuring MIG
MIG strategies, profiles, creating instances and targeting one MIG device.
Key points
MIG Manager is the GPU Operator component that applies MIG layouts. It watches the node label and reconfigures the GPUs.
What NVIDIA says (2)
“nvidia.com/mig.config=all-1g.10gb”
“GPU Operator deploys MIG Manager to manage MIG configuration on nodes in your Kubernetes cluster.”
The strategy tells the Operator how MIG devices are exposed. You choose it at install time, before you configure MIG.
What NVIDIA says (2)
“single : MIG mode is enabled on all GPUs on a node.”
“mixed : MIG mode is not enabled on all GPUs on a node.”
A GPU instance profile is a fixed slice size, such as 1g.5gb. -lgip lists profiles with free and total counts.
What NVIDIA says (1)
“$ nvidia-smi mig -lgip”
-cgi creates GPU instances. You can name a profile by ID (9) or by name (3g.20gb). -C also creates the matching compute instances.
What NVIDIA says (2)
“sudo nvidia-smi mig -cgi 9,3g.20gb -C”
“Successfully created GPU instance ID 2 on GPU 0 using profile MIG 3g.20gb (ID 9)”
Each MIG device has its own UUID. CUDA_VISIBLE_DEVICES limits an app to the devices you name.
What NVIDIA says (2)
“$ nvidia-smi -L GPU 0: A100-SXM4-40GB”
“$ CUDA_VISIBLE_DEVICES=MIG-c7384736-a75d-5afc-978f-d2f1294409fd ./BlackScholes &”
Key terms: Multi-Instance GPU
Try it: MIG Partitioning