Workload Management
23% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization
3.1 Inference on Kubernetes
Deploying NIM with the NIM Operator, caching models and autoscaling.
Key points
The NIM Operator is a Kubernetes operator for NIM. You describe the NIM deployment you want in a custom resource, and the Operator makes it so.
What NVIDIA says (1)
“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”
Models are large. The NIM Operator can pre-cache them on cluster storage. New pods then start from the cache instead of downloading again.
What NVIDIA says (2)
“One key benefit of using the NIM Operator is its ability to pre-cache models and datasets.”
“Models, their many available profiles, and training datasets are large and can take a long time to download.”
A NIMService is a custom resource. kubectl lists it like any other resource. -A means all namespaces.
What NVIDIA says (1)
“View the NIM services custom resources: $ kubectl get nimservices.apps.nvidia.com -A”
Horizontal Pod Autoscaling adds or removes pod replicas based on metrics. The NIM Operator docs list Prometheus as the prerequisite.
What NVIDIA says (1)
“Configuring Horizontal Pod Autoscaling # Prerequisites # Prometheus installed on your cluster.”
Key terms: NVIDIA NIM NIM Operator Horizontal Pod Autoscaling
Try it: Training vs Inference Serving
3.2 Inference with Run:ai
Run:ai inference workloads, their prerequisites and custom servers.
Key points
Knative is a Kubernetes add-on for serving and autoscaling request-driven workloads. Run:ai inference workloads depend on it.
What NVIDIA says (1)
“Make sure Knative is properly installed by your administrator.”
Every Run:ai workload belongs to a project. A project's quota is the GPU share it is guaranteed.
What NVIDIA says (1)
“The inference workload is assigned to a project and is affected by the project’s quota.”
The Custom server option lets you bring your own container image and server configuration.
What NVIDIA says (1)
“allows you to bring your own container image and server configuration - for example, when using an inference server not natively supported by NVIDIA Run:ai, such as SGLang.”
Run:ai inference covers small and very large models. Large LLMs can span several nodes.
What NVIDIA says (1)
“The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”
Key terms: Run:ai project NVIDIA NIM Knative
Try it: Training vs Inference Serving
3.3 Training with Slurm
Job scripts, GPU requests and containers through Pyxis and Enroot.
Key points
#SBATCH lines in a job script are Slurm options. --gpus sets how many GPUs the job gets.
What NVIDIA says (1)
“#SBATCH -p defq #assuming node in defq has a GPU #SBATCH --gpus=1”
A job script, also called a batch file, holds the Slurm options and the commands to run.
What NVIDIA says (1)
“it is usually more convenient to send jobs to Slurm with the sbatch command acting on a job script.”
Pyxis is a Slurm SPANK plugin. Enroot is the tool that turns container images into unprivileged sandboxes.
What NVIDIA says (1)
“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”
The job pulls a small image and prints its OS name. If it prints, Pyxis and Enroot are working.
What NVIDIA says (1)
“srun --container-image=ubuntu grep PRETTY /etc/os-release pyxis: importing docker image: ubuntu”
Try it: Slurm Scheduler
3.4 Training with Run:ai
Priority, preemption, distributed training and checkpoints.
Key points
Preemptible means the scheduler may pause it to free GPUs for more important work. This lets training use spare GPUs beyond the project's quota.
What NVIDIA says (1)
“By default, training workloads are assigned a Low priority and are preemptible .”
Distributed training spans nodes, so the workers must coordinate over the network.
What NVIDIA says (1)
“Multi-GPU training uses multiple GPUs within a single node, whereas distributed training spans multiple nodes and typically requires coordination between them.”
A checkpoint is a saved copy of training progress. A resumed workload may land on another node, so local disk may not be there.
What NVIDIA says (1)
“Always use shared network storage (e.g., NFS). When a preempted workload is resumed, it may be scheduled on a different node than before.”
You choose Workers & master or Workers only. The Master inherits the Worker setup unless you override it.
What NVIDIA says (1)
“By default, the Master uses the Worker configuration for shared fields.”
Key terms: Preemption Checkpoint
3.5 System management tools
DCGM and nvidia-smi for finding and clearing GPU problems.
Key points
DCGM (Data Center GPU Manager) diagnostics run in levels. Higher levels take longer and test more.
What NVIDIA says (1)
“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”
Discovery is the first check. If a GPU is missing here, deeper tests will not help.
What NVIDIA says (1)
“You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”
dmon is device monitoring. It prints a compact line per cycle.
What NVIDIA says (1)
“This tool allows the user to see one line of monitoring data per monitoring cycle.”
A GPU reset clears hardware and software state on the GPU. Use -i to target one GPU.
What NVIDIA says (1)
“Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”
Key terms: Data Center GPU Manager nvidia-smi
Try it: DCGM Monitoring
3.6 Sharing resources between teams
Quota, over quota, fair share and GPU fractions.
Key points
Over quota means using more than your guaranteed share when GPUs are free. Fairness means that share is returned when its owner needs it.
What NVIDIA says (1)
“To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”
Fair share balances scheduling priority by how much each account has used.
What NVIDIA says (1)
“Similar to the Slurm command sshare, an administrator can display the Slurm account hierarchy with the fairshare command in cmsh”
Fractions let many small workloads share one GPU, which raises utilization and lets quota be set more precisely.
What NVIDIA says (1)
“With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users”
Key terms: Deserved quota Preemption Fair share GPU fractions
3.7 Containers from NGC
Reading NGC image names, tags and the docker run options.
Key points
NGC (NVIDIA GPU Cloud) is NVIDIA's catalog of GPU-optimized software. nvcr.io is its container registry.
What NVIDIA says (1)
“nvcr.io : The name of the container registry, which for the NGC container registry is nvcr.io .”
--gpus all exposes the GPUs. -it is interactive. --rm deletes the container on exit. -v mounts a directory.
What NVIDIA says (1)
“A run command looks similar to: docker run --gpus all -it --rm -v local_dir:container_dir nvcr.io/nvidia/caffe2:<xx.xx>”
A tag names a specific image version. Always state the tag you want from the catalog.
What NVIDIA says (1)
“If you choose not to add a tag to an image, by default the word “latest” is added as the tag, however all NGC containers have an explicit version tag.”
Key terms: NGC
Try it: NGC Container Flow