Workload Management

23% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization

3.1 Inference on Kubernetes

Official objective: “Deploy inference workloads with Kubernetes”

Deploying NIM with the NIM Operator, caching models and autoscaling.

Key points

  1. The NIM Operator is a Kubernetes operator for NIM. You describe the NIM deployment you want in a custom resource, and the Operator makes it so.

    What NVIDIA says (1)

    “The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”

    — NVIDIA NIM Operator

  2. Models are large. The NIM Operator can pre-cache them on cluster storage. New pods then start from the cache instead of downloading again.

    What NVIDIA says (2)

    “One key benefit of using the NIM Operator is its ability to pre-cache models and datasets.”

    — NVIDIA NIM Operator

    “Models, their many available profiles, and training datasets are large and can take a long time to download.”

    — NVIDIA NIM Operator

  3. A NIMService is a custom resource. kubectl lists it like any other resource. -A means all namespaces.

    What NVIDIA says (1)

    “View the NIM services custom resources: $ kubectl get nimservices.apps.nvidia.com -A”

    — NVIDIA NIM Operator: NIM Service

  4. Horizontal Pod Autoscaling adds or removes pod replicas based on metrics. The NIM Operator docs list Prometheus as the prerequisite.

    What NVIDIA says (1)

    “Configuring Horizontal Pod Autoscaling # Prerequisites # Prometheus installed on your cluster.”

    — NVIDIA NIM Operator: NIM Service

Key terms: NVIDIA NIM NIM Operator Horizontal Pod Autoscaling

Try it: Training vs Inference Serving

Practice 3.1 (4 questions)

3.2 Inference with Run:ai

Official objective: “Deploy inference workloads with Run:ai”

Run:ai inference workloads, their prerequisites and custom servers.

Key points

  1. Knative is a Kubernetes add-on for serving and autoscaling request-driven workloads. Run:ai inference workloads depend on it.

    What NVIDIA says (1)

    “Make sure Knative is properly installed by your administrator.”

    — NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

  2. Every Run:ai workload belongs to a project. A project's quota is the GPU share it is guaranteed.

    What NVIDIA says (1)

    “The inference workload is assigned to a project and is affected by the project’s quota.”

    — NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

  3. The Custom server option lets you bring your own container image and server configuration.

    What NVIDIA says (1)

    “allows you to bring your own container image and server configuration - for example, when using an inference server not natively supported by NVIDIA Run:ai, such as SGLang.”

    — NVIDIA Run:ai: Deploy Inference Workloads with a Custom Server

  4. Run:ai inference covers small and very large models. Large LLMs can span several nodes.

    What NVIDIA says (1)

    “The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”

    — NVIDIA Run:ai Inference Overview

Key terms: Run:ai project NVIDIA NIM Knative

Try it: Training vs Inference Serving

Practice 3.2 (4 questions)

3.3 Training with Slurm

Official objective: “Deploy training workloads with Slurm”

Job scripts, GPU requests and containers through Pyxis and Enroot.

Key points

  1. #SBATCH lines in a job script are Slurm options. --gpus sets how many GPUs the job gets.

    What NVIDIA says (1)

    “#SBATCH -p defq #assuming node in defq has a GPU #SBATCH --gpus=1”

    — NVIDIA Base Command Manager 11 User Manual (PDF)

  2. A job script, also called a batch file, holds the Slurm options and the commands to run.

    What NVIDIA says (1)

    “it is usually more convenient to send jobs to Slurm with the sbatch command acting on a job script.”

    — NVIDIA Base Command Manager 11 User Manual (PDF)

  3. Pyxis is a Slurm SPANK plugin. Enroot is the tool that turns container images into unprivileged sandboxes.

    What NVIDIA says (1)

    “The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  4. The job pulls a small image and prints its OS name. If it prints, Pyxis and Enroot are working.

    What NVIDIA says (1)

    “srun --container-image=ubuntu grep PRETTY /etc/os-release pyxis: importing docker image: ubuntu”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Slurm Pyxis Enroot

Try it: Slurm Scheduler

Practice 3.3 (4 questions)

3.4 Training with Run:ai

Official objective: “Deploy training workloads with Run:ai”

Priority, preemption, distributed training and checkpoints.

Key points

  1. Preemptible means the scheduler may pause it to free GPUs for more important work. This lets training use spare GPUs beyond the project's quota.

    What NVIDIA says (1)

    “By default, training workloads are assigned a Low priority and are preemptible .”

    — NVIDIA Run:ai: Train Models Using a Standard Training Workload

  2. Distributed training spans nodes, so the workers must coordinate over the network.

    What NVIDIA says (1)

    “Multi-GPU training uses multiple GPUs within a single node, whereas distributed training spans multiple nodes and typically requires coordination between them.”

    — NVIDIA Run:ai: Distributed Training Workloads

  3. A checkpoint is a saved copy of training progress. A resumed workload may land on another node, so local disk may not be there.

    What NVIDIA says (1)

    “Always use shared network storage (e.g., NFS). When a preempted workload is resumed, it may be scheduled on a different node than before.”

    — NVIDIA Run:ai: Checkpointing Preemptible Training Workloads

  4. You choose Workers & master or Workers only. The Master inherits the Worker setup unless you override it.

    What NVIDIA says (1)

    “By default, the Master uses the Worker configuration for shared fields.”

    — NVIDIA Run:ai: Distributed Training Workloads

Key terms: Preemption Checkpoint

Practice 3.4 (4 questions)

3.5 System management tools

Official objective: “Use system management tools to troubleshoot issues”

DCGM and nvidia-smi for finding and clearing GPU problems.

Key points

  1. DCGM (Data Center GPU Manager) diagnostics run in levels. Higher levels take longer and test more.

    What NVIDIA says (1)

    “Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”

    — DCGM Diagnostics

  2. Discovery is the first check. If a GPU is missing here, deeper tests will not help.

    What NVIDIA says (1)

    “You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”

    — DCGM User Guide: Getting Started

  3. dmon is device monitoring. It prints a compact line per cycle.

    What NVIDIA says (1)

    “This tool allows the user to see one line of monitoring data per monitoring cycle.”

    — nvidia-smi documentation

  4. A GPU reset clears hardware and software state on the GPU. Use -i to target one GPU.

    What NVIDIA says (1)

    “Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”

    — nvidia-smi documentation

Key terms: Data Center GPU Manager nvidia-smi

Try it: DCGM Monitoring

Practice 3.5 (4 questions)

3.6 Sharing resources between teams

Official objective: “Allocate resources between teams with Run:ai, Slurm and Kubernetes”

Quota, over quota, fair share and GPU fractions.

Key points

  1. Over quota means using more than your guaranteed share when GPUs are free. Fairness means that share is returned when its owner needs it.

    What NVIDIA says (1)

    “To maintain fairness, the NVIDIA Run:ai Scheduler preempts workload a1 (1 GPU), freeing up resources for team-b.”

    — NVIDIA Run:ai: Over Quota, Fairness and Preemption

  2. Fair share balances scheduling priority by how much each account has used.

    What NVIDIA says (1)

    “Similar to the Slurm command sshare, an administrator can display the Slurm account hierarchy with the fairshare command in cmsh”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  3. Fractions let many small workloads share one GPU, which raises utilization and lets quota be set more precisely.

    What NVIDIA says (1)

    “With GPU fractions, you can divide the GPU/s memory into smaller chunks and share the GPU/s compute resources between different workloads and users”

    — NVIDIA Run:ai: GPU Fractions

Key terms: Deserved quota Preemption Fair share GPU fractions

Practice 3.6 (3 questions)

3.7 Containers from NGC

Official objective: “Deploy containers from NGC”

Reading NGC image names, tags and the docker run options.

Key points

  1. NGC (NVIDIA GPU Cloud) is NVIDIA's catalog of GPU-optimized software. nvcr.io is its container registry.

    What NVIDIA says (1)

    “nvcr.io : The name of the container registry, which for the NGC container registry is nvcr.io .”

    — NGC Catalog User Guide

  2. --gpus all exposes the GPUs. -it is interactive. --rm deletes the container on exit. -v mounts a directory.

    What NVIDIA says (1)

    “A run command looks similar to: docker run --gpus all -it --rm -v local_dir:container_dir nvcr.io/nvidia/caffe2:<xx.xx>”

    — NGC Catalog User Guide

  3. A tag names a specific image version. Always state the tag you want from the catalog.

    What NVIDIA says (1)

    “If you choose not to add a tag to an image, by default the word “latest” is added as the tag, however all NGC containers have an explicit version tag.”

    — NGC Catalog User Guide

Key terms: NGC

Try it: NGC Container Flow

Practice 3.7 (3 questions)