NCP-AII hands-on labs

Guided labs in the Aegis lab simulator. They use simulated, illustrative command output, not real hardware. Each lab is tagged to official objectives and cites the NVIDIA source for the concept it practises.

MIG Partitioning: simulator lab

Objectives 2.2

Enable MIG mode and split a GPU into GPU instances and compute instances.

Open the lab simulator Pick “MIG Partitioning” from the lab list.

Try this

  • Enable MIG mode on one GPU.
  • Create GPU instances, then their compute instances.
What NVIDIA says (2)

“By default, MIG mode is not enabled on the GPU.”

— MIG User Guide: Getting Started with MIG

“Once the GPU instances are created, you need to create the corresponding Compute Instances (CI).”

— MIG User Guide: Getting Started with MIG

AllReduce Deep Dive: simulator lab

Objectives 4.10

See how NCCL runs all-reduce across nodes and read busbw correctly.

Open the lab simulator Pick “AllReduce Deep Dive” from the lab list.

Try this

  • Run a multi-node all_reduce test.
  • Read the busbw column.
What NVIDIA says (2)

“Run 64 MPI processes on nodes with 8 GPUs each, for a total of 64 GPUs spread across 8 nodes.”

— NVIDIA nccl-tests

“See the Performance page for explanation about numbers, and in particular the "busbw" column.”

— NVIDIA nccl-tests

InfiniBand Fabric: simulator lab

Objectives 4.4, 4.5

Check an InfiniBand port and its error counters, then localize a bad link with ibdiagnet.

Open the lab simulator Pick “InfiniBand Fabric” from the lab list.

Try this

  • Read a port's state and error counters.
  • Scan the fabric with ibdiagnet and find the bad link.
What NVIDIA says (2)

“Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Link width and speed checks”

— MLNX_OFED: InfiniBand Fabric Utilities

GPUDirect Storage: simulator lab

Objectives 1.11, 4.14

Compare the CPU bounce-buffer path with the GDS direct path, check GDS support and measure both with gdsio.

Open the lab simulator Pick “GPUDirect Storage” from the lab list.

Try this

  • Verify the platform with gdscheck.
  • Run gdsio on the GDS path.
What NVIDIA says (2)

“using gdscheck you should see the following output if the IOMMU is disabled on the system:”

— GPUDirect Storage Best Practices Guide

“-x 0 , the IO data path, in this case GDS.”

— GPUDirect Storage Configuration and Benchmarking Guide

ECC Error Lifecycle: simulator lab

Objectives 5.1

Follow GPU memory errors from a clean baseline to an uncorrectable Xid 48, then contain the node.

Open the lab simulator Pick “ECC Error Lifecycle” from the lab list.

Try this

  • Watch ECC error counts with DCGM.
  • Pick the documented action for Xid 48.
What NVIDIA says (2)

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— Xid Errors: Analyzing the Xid Catalog

“A GPU reset or node reboot is needed to clear this error.”

— Xid Errors: Analyzing the Xid Catalog

Slurm Scheduler: simulator lab

Objectives 3.3, 5.1

Run GPU jobs under Slurm with NVIDIA's container plugin, then drain and resume a node.

Open the lab simulator Pick “Slurm Scheduler” from the lab list.

Try this

  • Run a containerized GPU job through srun.
  • Drain a suspect node, then resume it.
What NVIDIA says (2)

“Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”

— NVIDIA/pyxis: Container plugin for Slurm

“If an issue is found, it should be removed from the batch partition for initial triage.”

— DGX SuperPOD Administration Guide: System Health Checks and Debugging

NGC Container Flow: simulator lab

Objectives 3.6

Run a GPU container and confirm the GPUs are visible inside it.

Open the lab simulator Pick “NGC Container Flow” from the lab list.

Try this

  • Run a container with GPU access.
  • Check nvidia-smi output inside the container.
What NVIDIA says (2)

“sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”

— NVIDIA Container Toolkit: Running a Sample Workload

“you can verify your installation by running a sample workload.”

— NVIDIA Container Toolkit: Running a Sample Workload

CUDA Stack Verification: simulator lab

Objectives 3.4

Verify the installed NVIDIA driver on a node.

Open the lab simulator Pick “CUDA Stack Verification” from the lab list.

Try this

  • Read the loaded driver version.
  • Confirm it matches the planned version.
What NVIDIA says (2)

“When the driver is loaded, the driver version can be found by executing the following command: $ cat /proc/driver/nvidia/version”

— NVIDIA Driver Installation Guide: Post-installation Actions

“it is best to manually ensure the correct version of the kernel headers and development packages are installed prior to installing the NVIDIA driver”

— NVIDIA Driver Installation Guide: Pre-installation Actions