NCP-AIO hands-on labs

Guided labs in the Aegis lab simulator. They use simulated, illustrative command output, not real hardware. Each lab is tagged to official objectives and cites the NVIDIA source for the concept it practises.

Slurm Scheduler: simulator lab

Objectives 2.1, 3.3

Run GPU jobs under Slurm with the Pyxis container plugin, then drain and resume a node.

Open the lab simulator Pick “Slurm Scheduler” from the lab list.

Try this

  • Run a containerized GPU job through srun.
  • Drain a suspect node, then resume it.
What NVIDIA says (2)

“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”

— NVIDIA Base Command Manager 11 Administrator Manual (PDF)

“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”

— NVIDIA Mission Control Administration Guide: Slurm Workload Management

Kubernetes GPU Ops: simulator lab

Objectives 2.4, 4.6

Check the GPU Operator pods, request GPUs in a pod, debug a Pending pod and spot time-slicing oversubscription.

Open the lab simulator Pick “Kubernetes GPU Ops” from the lab list.

Try this

  • Check that GPU Operator pods are running.
  • Read why a pod is Pending.
What NVIDIA says (2)

“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”

— GPU Operator: Installing the NVIDIA GPU Operator

“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”

— Time-Slicing GPUs in Kubernetes

MIG Partitioning: simulator lab

Objectives 2.5

Enable MIG mode and split a GPU into GPU instances and compute instances.

Open the lab simulator Pick “MIG Partitioning” from the lab list.

Try this

  • List the GPU instance profiles.
  • Create GPU instances with their compute instances.
What NVIDIA says (2)

“$ nvidia-smi mig -lgip”

— MIG User Guide: Getting Started with MIG

“Successfully created GPU instance ID 2 on GPU 0 using profile MIG 3g.20gb (ID 9)”

— MIG User Guide: Getting Started with MIG

NGC Container Flow: simulator lab

Objectives 3.7, 4.1

Pull a container from the NGC catalog and run it with GPU access.

Open the lab simulator Pick “NGC Container Flow” from the lab list.

Try this

  • Pull a tagged NGC image.
  • Run it with --gpus and confirm the GPUs inside.
What NVIDIA says (2)

“A run command looks similar to: docker run --gpus all -it --rm -v local_dir:container_dir nvcr.io/nvidia/caffe2:<xx.xx>”

— NGC Catalog User Guide

“If you choose not to add a tag to an image, by default the word “latest” is added as the tag, however all NGC containers have an explicit version tag.”

— NGC Catalog User Guide

DCGM Monitoring: simulator lab

Objectives 3.5

Watch live GPU telemetry with DCGM and set an alert on memory errors.

Open the lab simulator Pick “DCGM Monitoring” from the lab list.

Try this

  • Stream GPU metrics.
  • Write an alert on double-bit ECC errors.
What NVIDIA says (2)

“You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”

— DCGM User Guide: Getting Started

“Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”

— nvidia-smi documentation

Training vs Inference Serving: simulator lab

Objectives 3.1, 3.2

Compare a training job with an inference service and trade latency against throughput.

Open the lab simulator Pick “Training vs Inference Serving” from the lab list.

Try this

  • Profile an inference request.
  • Change batch size and watch latency and throughput.
What NVIDIA says (2)

“The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”

— NVIDIA Run:ai Inference Overview

“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”

— NVIDIA NIM Operator

AllReduce Deep Dive: simulator lab

Objectives 4.4

See how NCCL runs all-reduce across GPUs and nodes, and read the result.

Open the lab simulator Pick “AllReduce Deep Dive” from the lab list.

Try this

  • Run an all-reduce test.
  • Read the bandwidth result.
What NVIDIA says (2)

“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”

— NCCL Environment Variables

“The NCCL_P2P_DISABLE variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”

— NCCL Environment Variables

GPUDirect Storage: simulator lab

Objectives 4.4, 4.5

Compare the CPU bounce-buffer path with the GPUDirect Storage path, check support and benchmark both.

Open the lab simulator Pick “GPUDirect Storage” from the lab list.

Try this

  • Check the platform with gdscheck.
  • Benchmark with gdsio.
What NVIDIA says (2)

“In the /etc/cufile.json file, verify that allow_compat_mode is set to true . gdscheck -p displays whether the allow_compat_mode property is set to true .”

— GPUDirect Storage Troubleshooting Guide

“the number of processes/threads generating IO is critical to determining maximum performance. The gdsio tool provides for specifying random reads or random writes”

— GPUDirect Storage Configuration and Benchmarking Guide

Storage Bottleneck: simulator lab

Objectives 4.5

Find a storage bottleneck that starves GPUs and move the dataset to faster storage.

Open the lab simulator Pick “Storage Bottleneck” from the lab list.

Try this

  • Spot the I/O bottleneck in the simulated metrics.
  • Stage the data and compare.
What NVIDIA says (2)

“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”

— NVIDIA Mission Control Administration Guide: Overview

“The optimal path between a GPU and NIC will be one PCIe switch path designated as PIX. The least optimal path is designated as SYS”

— GPUDirect Storage Configuration and Benchmarking Guide