NCP-AII hands-on labs
Guided labs in the Aegis lab simulator. They use simulated, illustrative command output, not real hardware. Each lab is tagged to official objectives and cites the NVIDIA source for the concept it practises.
MIG Partitioning: simulator lab
Enable MIG mode and split a GPU into GPU instances and compute instances.
Try this
- Enable MIG mode on one GPU.
- Create GPU instances, then their compute instances.
What NVIDIA says (2)
“By default, MIG mode is not enabled on the GPU.”
“Once the GPU instances are created, you need to create the corresponding Compute Instances (CI).”
NVLink Topology: simulator lab
Read the NVLink topology of an 8-GPU node, check link error counters and run a single-node all-reduce.
Try this
- Print the GPU topology matrix.
- Run all_reduce_perf on one node and read the result.
What NVIDIA says (2)
“Topology connections and affinities matrix between the GPUs and NICs in the system nvidia-smi topo -m”
“These tests check both the performance and the correctness of NCCL operations.”
AllReduce Deep Dive: simulator lab
See how NCCL runs all-reduce across nodes and read busbw correctly.
Try this
- Run a multi-node all_reduce test.
- Read the busbw column.
What NVIDIA says (2)
“Run 64 MPI processes on nodes with 8 GPUs each, for a total of 64 GPUs spread across 8 nodes.”
“See the Performance page for explanation about numbers, and in particular the "busbw" column.”
InfiniBand Fabric: simulator lab
Check an InfiniBand port and its error counters, then localize a bad link with ibdiagnet.
Try this
- Read a port's state and error counters.
- Scan the fabric with ibdiagnet and find the bad link.
What NVIDIA says (2)
“Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”
“Link width and speed checks”
GPUDirect Storage: simulator lab
Compare the CPU bounce-buffer path with the GDS direct path, check GDS support and measure both with gdsio.
Try this
- Verify the platform with gdscheck.
- Run gdsio on the GDS path.
What NVIDIA says (2)
“using gdscheck you should see the following output if the IOMMU is disabled on the system:”
“-x 0 , the IO data path, in this case GDS.”
ECC Error Lifecycle: simulator lab
Follow GPU memory errors from a clean baseline to an uncorrectable Xid 48, then contain the node.
Try this
- Watch ECC error counts with DCGM.
- Pick the documented action for Xid 48.
What NVIDIA says (2)
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“A GPU reset or node reboot is needed to clear this error.”
XID Fault Drill: simulator lab
Respond to three Xid errors (48, 79, 74) and pick the right next action for each.
Try this
- Read each Xid in the kernel log.
- Match it to the catalog and pick the action.
What NVIDIA says (2)
“This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”
“This event is often caused by hardware failures on the PCI Express link”
Slurm Scheduler: simulator lab
Run GPU jobs under Slurm with NVIDIA's container plugin, then drain and resume a node.
Try this
- Run a containerized GPU job through srun.
- Drain a suspect node, then resume it.
What NVIDIA says (2)
“Pyxis is a SPANK plugin for the Slurm Workload Manager. It allows unprivileged cluster users to run containerized tasks through the srun command.”
“If an issue is found, it should be removed from the batch partition for initial triage.”
NGC Container Flow: simulator lab
Run a GPU container and confirm the GPUs are visible inside it.
Try this
- Run a container with GPU access.
- Check nvidia-smi output inside the container.
What NVIDIA says (2)
“sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”
“you can verify your installation by running a sample workload.”
CUDA Stack Verification: simulator lab
Verify the installed NVIDIA driver on a node.
Try this
- Read the loaded driver version.
- Confirm it matches the planned version.
What NVIDIA says (2)
“When the driver is loaded, the driver version can be found by executing the following command: $ cat /proc/driver/nvidia/version”
“it is best to manually ensure the correct version of the kernel headers and development packages are installed prior to installing the NVIDIA driver”