NCP-AIO hands-on labs
Guided labs in the Aegis lab simulator. They use simulated, illustrative command output, not real hardware. Each lab is tagged to official objectives and cites the NVIDIA source for the concept it practises.
Slurm Scheduler: simulator lab
Run GPU jobs under Slurm with the Pyxis container plugin, then drain and resume a node.
Try this
- Run a containerized GPU job through srun.
- Drain a suspect node, then resume it.
What NVIDIA says (2)
“The Pyxis plugin requires the Enroot utility, and allows the user’s jobs to be executed seamlessly over Enroot in unprivileged containers.”
“scontrol update nodename = a06-p1-dgx-02-c [ 01 -04,06-11,13-14,16-18 ] state = drain reason = "maintenance"”
Kubernetes GPU Ops: simulator lab
Check the GPU Operator pods, request GPUs in a pod, debug a Pending pod and spot time-slicing oversubscription.
Try this
- Check that GPU Operator pods are running.
- Read why a pod is Pending.
What NVIDIA says (2)
“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”
MIG Partitioning: simulator lab
Enable MIG mode and split a GPU into GPU instances and compute instances.
Try this
- List the GPU instance profiles.
- Create GPU instances with their compute instances.
What NVIDIA says (2)
“$ nvidia-smi mig -lgip”
“Successfully created GPU instance ID 2 on GPU 0 using profile MIG 3g.20gb (ID 9)”
NGC Container Flow: simulator lab
Pull a container from the NGC catalog and run it with GPU access.
Try this
- Pull a tagged NGC image.
- Run it with --gpus and confirm the GPUs inside.
What NVIDIA says (2)
“A run command looks similar to: docker run --gpus all -it --rm -v local_dir:container_dir nvcr.io/nvidia/caffe2:<xx.xx>”
“If you choose not to add a tag to an image, by default the word “latest” is added as the tag, however all NGC containers have an explicit version tag.”
DCGM Monitoring: simulator lab
Watch live GPU telemetry with DCGM and set an alert on memory errors.
Try this
- Stream GPU metrics.
- Write an alert on double-bit ECC errors.
What NVIDIA says (2)
“You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”
“Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”
Training vs Inference Serving: simulator lab
Compare a training job with an inference service and trade latency against throughput.
Try this
- Profile an inference request.
- Change batch size and watch latency and throughput.
What NVIDIA says (2)
“The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”
“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”
AllReduce Deep Dive: simulator lab
See how NCCL runs all-reduce across GPUs and nodes, and read the result.
Try this
- Run an all-reduce test.
- Read the bandwidth result.
What NVIDIA says (2)
“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”
“The NCCL_P2P_DISABLE variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”
GPUDirect Storage: simulator lab
Compare the CPU bounce-buffer path with the GPUDirect Storage path, check support and benchmark both.
Try this
- Check the platform with gdscheck.
- Benchmark with gdsio.
What NVIDIA says (2)
“In the /etc/cufile.json file, verify that allow_compat_mode is set to true . gdscheck -p displays whether the allow_compat_mode property is set to true .”
“the number of processes/threads generating IO is critical to determining maximum performance. The gdsio tool provides for specifying random reads or random writes”
Storage Bottleneck: simulator lab
Find a storage bottleneck that starves GPUs and move the dataset to faster storage.
Try this
- Spot the I/O bottleneck in the simulated metrics.
- Stage the data and compare.
What NVIDIA says (2)
“High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”
“The optimal path between a GPU and NIC will be one PCIe switch path designated as PIX. The least optimal path is designated as SYS”