NCA-AIIO hands-on labs

Guided labs in the Aegis lab simulator. They use realistic simulated command output, not real hardware. Each lab is tagged to official objectives and cites the NVIDIA source for the concept it practises.

MIG Partitioning: simulator lab

Objectives 3.4

Enable Multi-Instance GPU and create isolated GPU instances.

Open the lab simulator Pick “MIG Partitioning” from the lab list.

Try this

  • Enable MIG mode.
  • Create instances and list them.
What NVIDIA says (9)

“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”

— MIG User Guide: Introduction

“By default, MIG mode is not enabled on the GPU.”

— MIG User Guide: Getting Started with MIG

“Without creating GPU instances (and corresponding compute instances), CUDA workloads cannot be run on the GPU.”

— MIG User Guide: Getting Started with MIG

“MIG 1g.10gb 1/8 1/7 1 NVDEC /1 JPEG /0 OFA 1/8 1 7”

— MIG User Guide: Supported MIG Profiles

“Once the GPU instances are created, you need to create the corresponding Compute Instances (CI).”

— MIG User Guide: Getting Started with MIG

“MIG 3g.20gb Device 0: (UUID: MIG-c7384736-a75d-5afc-978f-d2f1294409fd)”

— MIG User Guide: Getting Started with MIG

“CUDA_VISIBLE_DEVICES=MIG-c7384736-a75d-5afc-978f-d2f1294409fd ./BlackScholes”

— MIG User Guide: Getting Started with MIG

“sudo nvidia-smi mig -dci && sudo nvidia-smi mig -dgi”

— MIG User Guide: Getting Started with MIG

“the created MIG devices are not persistent across system reboots”

— MIG User Guide: Getting Started with MIG

ECC Error Lifecycle: simulator lab

Objectives 3.1, 3.3

Follow an ECC error from DCGM counters to the Xid 48 recovery workflow.

Open the lab simulator Pick “ECC Error Lifecycle” from the lab list.

Try this

  • Read field 311 in DCGM.
  • Pick the recovery action for Xid 48.
What NVIDIA says (8)

“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”

— DCGM Field IDs

“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”

— NVIDIA Xid Catalog

“DCGM_FI_DEV_ECC_SBE_VOL_TOTAL 310 Total single bit volatile ECC errors”

— DCGM Field IDs

“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors”

— DCGM Field IDs

“On GPUs that support row remapping, starting with NVIDIA® Ampere archtecture GPUs, these events provide details on row remapper activity.”

— NVIDIA Xid Catalog

“48 ROBUST_CHANNEL_GPU_ECC_DBE Double Bit ECC Error”

— NVIDIA Xid Catalog

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU. This is also reported back to the user application. A GPU reset or node reboot is needed to clear this error.”

— NVIDIA Xid Catalog

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— NVIDIA Xid Catalog

CUDA Stack Verification: simulator lab

Objectives 1.1

Check that the driver and CUDA versions on a node can run a given application, using CUDA compatibility rules.

Open the lab simulator Pick “CUDA Stack Verification” from the lab list.

Try this

  • Read the driver and CUDA versions from simulated nvidia-smi output.
  • Decide whether forward compatibility is needed.
What NVIDIA says (8)

“CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”

— NVIDIA CUDA Compatibility

“NVIDIA-SMI 535.86.10 Driver Version: 535.86.10 CUDA Version: 12.2”

— NVIDIA Container Toolkit: Running a Sample Workload

“Backwards compatibility ensures that a newer NVIDIA driver can be used with an older CUDA Toolkit.”

— CUDA Compatibility: Why CUDA Compatibility

“Minor version and forward compatibility ensure that an older NVIDIA driver can be used with a newer CUDA Toolkit.”

— CUDA Compatibility: Why CUDA Compatibility

“Drivers have always been backwards compatible with CUDA.”

— CUDA Compatibility: FAQ

“It’s mainly intended to support applications built on newer CUDA Toolkits to run on systems installed with an older NVIDIA Linux GPU driver from different major release families.”

— CUDA Compatibility: Forward Compatibility

“NVIDIA creates an updated set of Docker containers for the frameworks monthly.”

— NVIDIA Deep Learning Frameworks User Guide (NGC containers)

“where deep learning frameworks are tuned, optimized, tested, and containerized for your use.”

— NVIDIA Deep Learning Frameworks User Guide (NGC containers)

NGC Container Flow: simulator lab

Objectives 1.1

Pull a GPU-optimized container from the NGC catalog and run it with the NVIDIA Container Toolkit.

Open the lab simulator Pick “NGC Container Flow” from the lab list.

Try this

  • Pull a framework container.
  • Run it with GPU access and confirm the GPUs are visible.
What NVIDIA says (6)

“The NGC Catalog consists of containers, pretrained models, Helm charts for Kubernetes deployments, and industry-specific AI toolkits with software development kits (SDKs).”

— NGC Catalog User Guide

“It currently includes: The NVIDIA Container Runtime ( nvidia-container-runtime ) The NVIDIA Container Toolkit CLI ( nvidia-ctk )”

— NVIDIA Container Toolkit

“sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”

— NVIDIA Container Toolkit: Running a Sample Workload

“As of Docker release 19.03, NVIDIA GPUs are natively supported as devices in the Docker runtime.”

— NVIDIA Deep Learning Frameworks User Guide (NGC containers)

“docker run --gpus all -it --rm –v local_dir:container_dir nvcr.io/nvidia/tensorflow:<xx.xx>-tf2-py3”

— NVIDIA Deep Learning Frameworks User Guide (NGC containers)

“nvidia-smi dmon”

— nvidia-smi documentation

Distributed Training (DDP): simulator lab

Objectives 1.2

Run a simulated multi-GPU training job and watch weights being adjusted over many iterations.

Open the lab simulator Pick “Distributed Training (DDP)” from the lab list.

Try this

  • Start a distributed training run.
  • Compare throughput as GPUs are added.
What NVIDIA says (9)

“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”

— What's the Difference Between Deep Learning Training and Inference?

“Data Parallelism (DP) replicates the model across multiple GPUs.”

— NVIDIA Megatron Bridge: Parallelisms Guide

“is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”

— Overview of NCCL — NCCL 2.32.3 documentation

“NCCL implements both collective communication and point-to-point send/receive primitives.”

— Overview of NCCL — NCCL 2.32.3 documentation

“Data batches are evenly distributed between GPUs and the data-parallel GPUs process them independently.”

— NVIDIA Megatron Bridge: Parallelisms Guide

“it sums the gradients of all model copies using all-reduce communication collectives.”

— NVIDIA Megatron Bridge: Parallelisms Guide

“The AllReduce operation performs reductions on data (for example, sum, min, max) across devices and stores the result in the receive buffer of every rank.”

— NCCL User Guide: Collective Operations

“Distributed Data Parallelism (DDP) keeps the model copies consistent by synchronizing parameter gradients across data-parallel GPUs before each parameter update.”

— NVIDIA Megatron Bridge: Parallelisms Guide

“bottleneck, limiting the performance and scalability of training and inference.”

— NVIDIA DALI User Guide

AllReduce Deep Dive: simulator lab

Objectives 2.7

Run an all-reduce across GPUs and nodes and see how NCCL uses NVLink inside a node and the network between nodes.

Open the lab simulator Pick “AllReduce Deep Dive” from the lab list.

Try this

  • Run all-reduce on one node, then two.
  • Compare bandwidth for each path.
What NVIDIA says (11)

“NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”

— Overview of NCCL — NCCL 2.32.3 documentation

“NCCL implements both collective communication and point-to-point send/receive primitives.”

— Overview of NCCL — NCCL 2.32.3 documentation

“It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”

— Overview of NCCL — NCCL 2.32.3 documentation

“INFO - Prints debug information.”

— NCCL Environment Variables

“A ring would do that operation in an order which follows the ring”

— nccl-tests: Performance reported by NCCL tests

“variable defines which algorithms NCCL will use.”

— NCCL Environment Variables

“we need 2(n-1) data transfers (x number of elements) to perform an allReduce operation.”

— nccl-tests: Performance reported by NCCL tests

“B = S/t * (2*(n-1)/n) = algbw * (2*(n-1)/n)”

— nccl-tests: Performance reported by NCCL tests

“Using this bus bandwidth, we can compare it with the hardware peak bandwidth, independently of the number of ranks used.”

— nccl-tests: Performance reported by NCCL tests

“The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”

— NCCL Environment Variables

“NCCL will instead fall back to another available transport such as IP sockets.”

— NCCL Environment Variables

InfiniBand Fabric: simulator lab

Objectives 2.8, 2.9

Check InfiniBand port state and fabric health, including the Subnet Manager.

Open the lab simulator Pick “InfiniBand Fabric” from the lab list.

Try this

  • Read simulated ibstat output.
  • Find the port that is down.
What NVIDIA says (10)

“InfiniBand is a high-performance, low latency, RDMA capable networking technology”

— DGX SuperPOD H100 Reference Architecture: Components

“All InfiniBand-compliant ULPs require a proper operation of a Subnet Manager (SM) running on the InfiniBand fabric, at all times.”

— MLNX_OFED Documentation: Introduction

“ibstat is a binary which displays basic information obtained from the local IB driver. Output includes LID, SMLID, port state, link width active, and port physical state.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Queries InfiniBand ports’ performance and error counters.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Calculates the BW of RDMA write between a pair of machines. One acts as a server and the other as a client.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Output includes LID, SMLID, port state, link width active, and port physical state.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Links in INIT state and unresponsive links detection Counters fetch Error counters check Routing checks Link width and speed checks”

— MLNX_OFED: InfiniBand Fabric Utilities

“Link width and speed checks”

— MLNX_OFED: InfiniBand Fabric Utilities

“ibstat is a binary which displays basic information obtained from the local IB driver.”

— MLNX_OFED: InfiniBand Fabric Utilities

RoCEv2 + PFC/ECN: simulator lab

Objectives 2.8

Configure lossless RoCE: MTU, Priority Flow Control (PFC) and ECN, then diagnose a PFC pause storm.

Open the lab simulator Pick “RoCEv2 + PFC/ECN” from the lab list.

Try this

  • Enable PFC on the RoCE priority.
  • Check ECN.
  • Read pause counters during a storm.
What NVIDIA says (13)

“the BlueField-3 SuperNIC provides best-in-class remote direct-memory access over converged Ethernet (RoCE) network connectivity between GPU servers”

— NVIDIA BlueField-3 Networking Platform User Guide: Introduction

“The regular Ethernet MTU applies on the RoCE frame.”

— MLNX_OFED: RDMA over Converged Ethernet (RoCE)

“RDMA over Converged Ethernet (RoCE) is a mechanism to provide this efficient data transfer with very low latencies on lossless Ethernet networks.”

— MLNX_OFED: RDMA over Converged Ethernet (RoCE)

“In order to function reliably, RoCE requires a form of flow control.”

— MLNX_OFED: RDMA over Converged Ethernet (RoCE)

“The normal and optimal way to use RoCE is to use Priority Flow Control (PFC). To use PFC, it must be enabled on all endpoints and switches in the flow path.”

— MLNX_OFED: RDMA over Converged Ethernet (RoCE)

“For example, PFC can provide lossless service for the RoCE traffic and best-effort service for the standard Ethernet traffic.”

— MLNX_OFED: Flow Control (PFC)

“It allows reliable communication by notifying all ends of communication when congestion occurs. This is done without dropping packets.”

— MLNX_OFED: Explicit Congestion Notification (ECN)

“cat /sys/class/net/<interface>/ecn/<protocol>/enable/X”

— MLNX_OFED: Explicit Congestion Notification (ECN)

“Calculates the BW of RDMA write between a pair of machines.”

— MLNX_OFED: InfiniBand Fabric Utilities

“Several ingress and egress counters per priority are supported. Run ethtool -S to get the full list of port counters.”

— MLNX_OFED: Flow Control (PFC)

“prio4_tx_pause: 26832”

— MLNX_OFED: Flow Control (PFC)

“PFC storm prevention enables toggling between default and auto modes.”

— MLNX_OFED: Flow Control (PFC)

“The stall prevention timeout is configured to 8 seconds by default. Auto mode sets the stall prevention timeout to be 100 msec.”

— MLNX_OFED: Flow Control (PFC)

NCCL Fallback Drill: simulator lab

Objectives 2.7

Find out why NCCL fell back from InfiniBand to IP sockets, and fix it.

Open the lab simulator Pick “NCCL Fallback Drill” from the lab list.

Try this

  • Read the NCCL debug output.
  • Find the setting that disabled the IB/RoCE transport.
What NVIDIA says (7)

“The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”

— NCCL Environment Variables

“NCCL will instead fall back to another available transport such as IP sockets.”

— NCCL Environment Variables

“It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”

— Overview of NCCL — NCCL 2.32.3 documentation

“INFO - Prints debug information.”

— NCCL Environment Variables

“ibstat is a binary which displays basic information obtained from the local IB driver.”

— MLNX_OFED: InfiniBand Fabric Utilities

“The NCCL_IB_HCA variable specifies which Host Channel Adapter (RDMA) interfaces to use for communication.”

— NCCL Environment Variables

“Using this bus bandwidth, we can compare it with the hardware peak bandwidth, independently of the number of ranks used.”

— nccl-tests: Performance reported by NCCL tests

Storage Bottleneck: simulator lab

Objectives 2.5

Find a storage bottleneck that starves GPUs, cache the dataset on local NVMe, and offload preprocessing with DALI.

Open the lab simulator Pick “Storage Bottleneck” from the lab list.

Try this

  • Spot the I/O bottleneck in the simulated metrics.
  • Stage data to local NVMe and compare.
What NVIDIA says (8)

“The key I/O operation in DL training is re-read.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

“In addition, the DGX H100 system provides local NVMe storage that can also be used for caching or staging data.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

“bottleneck, limiting the performance and scalability of training and inference.”

— NVIDIA DALI User Guide

“It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

“Ideally, data is cached during the first read of the dataset, so data does not have to be retrieved across the network.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

“Reading files from cache can be an order of magnitude faster than from remote storage.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

“DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the”

— NVIDIA DALI User Guide

“NVIDIA GPUDirect Storage® (GDS) provides a way to read data from the remote filesystem or local NVMe directly into GPU memory providing higher sustained I/O performance”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

GPUDirect Storage: simulator lab

Objectives 2.5

Compare a normal file read through CPU memory with a GPUDirect Storage read straight into GPU memory.

Open the lab simulator Pick “GPUDirect Storage” from the lab list.

Try this

  • Run both read paths.
  • Compare CPU load and throughput.
What NVIDIA says (11)

“GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which avoids a bounce buffer through the CPU.”

— GPUDirect Storage Overview

“an extra copy through a bounce buffer in the CPU is necessary, which introduces latency and lowers effective bandwidth.”

— GPUDirect Storage Overview

“GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which”

— GPUDirect Storage Overview

“The direct data path that GDS provides relies on the availability of file system drivers that are enabled with”

— GPUDirect Storage Overview

“/usr/local/cuda-<x>.<y>/gds/tools/gdscheck.py -p”

— NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

“Creating this direct path involves distributed file systems such as NFSoRDMA, DDN EXAScaler parallel file system solutions (based on the Lustre file system), Amazon FSx for Lustre, and WekaFS”

— NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

“compatibility mode is available for unsupported configurations that maps IO operations to a fallback path.”

— GPUDirect Storage Overview

“tool included when GDS is installed, gdsio, is covered and its use demonstrated.”

— GPUDirect Storage Configuration and Benchmarking Guide

“xfer_type : 0 - Storage -> GPU ( GDS ) 1 - Storage -> CPU 2 - Storage -> CPU -> GPU”

— GPUDirect Storage Configuration and Benchmarking Guide

“xfer_type : 0 - Storage -> GPU ( GDS )”

— GPUDirect Storage Configuration and Benchmarking Guide

“avoids a bounce buffer through the CPU. Using this direct path can relieve system bandwidth bottlenecks and decrease”

— GPUDirect Storage Overview

DCGM Monitoring: simulator lab

Objectives 3.1, 3.3

Collect GPU metrics with DCGM Exporter and Prometheus, and set an alert.

Open the lab simulator Pick “DCGM Monitoring” from the lab list.

Try this

  • Scrape the /metrics endpoint.
  • Write an alert on double-bit ECC errors.
What NVIDIA says (10)

“DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”

— NVIDIA DCGM Exporter

“Address of listening http server. Default: “:9400””

— NVIDIA DCGM Exporter

“curl localhost:9400/metrics”

— NVIDIA DCGM Exporter

“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors”

— DCGM Field IDs

“exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”

— NVIDIA DCGM Exporter

“serviceMonitorSelectorNilUsesHelmValues: false”

— GPU Telemetry: Setting up Prometheus and Grafana

“To add a dashboard for DCGM, you can use a standard dashboard that NVIDIA has made available, which can also be customized.”

— GPU Telemetry: Setting up Prometheus and Grafana

“https://grafana.com/grafana/dashboards/12239”

— GPU Telemetry: Setting up Prometheus and Grafana

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— NVIDIA Xid Catalog

“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”

— NVIDIA Xid Catalog

Slurm Scheduler: simulator lab

Objectives 3.2

Submit and inspect GPU jobs on a Slurm cluster managed with Base Command Manager.

Open the lab simulator Pick “Slurm Scheduler” from the lab list.

Try this

  • Submit a containerized GPU job with Pyxis.
  • Validate nodes with an NCCL test.
  • Drain and resume a node.
What NVIDIA says (10)

“NVIDIA Base Command Manager streamlines cluster provisioning, workload management, and infrastructure monitoring.”

— NVIDIA Base Command Manager Documentation

“Integrate with tools, like Slurm or NVIDIA Run:ai , to support traditional HPC or AI and analytics workloads across bare-metal and containerized environments.”

— NVIDIA Base Command Manager

“Slurm is a classic workload manager used to orchestrate complex workloads in a multi-node, batch-style, compute environment”

— DGX SuperPOD H100 Reference Architecture: Components

“#SBATCH --container-image nvcr.io\#nvidia/pytorch:21.12-py3”

— NVIDIA Pyxis: container plugin for Slurm

“the srun command will be in the queue until the nodes become available”

— NVIDIA DeepOps: Slurm Deployment Guide

“--container-image=[USER@][REGISTRY#]IMAGE[:TAG]|PATH”

— NVIDIA Pyxis: container plugin for Slurm

“srun_exports: NCCL_DEBUG=INFO”

— NVIDIA DeepOps: Slurm Deployment Guide

“The validation playbook will verify that Pyxis and Enroot can run GPU jobs”

— NVIDIA DeepOps: Slurm Deployment Guide

“Nodes which fail this check will be automatically drained in Slurm to prevent jobs run”

— NVIDIA DeepOps: Slurm Deployment Guide

“This tool will run periodically on idle nodes to validate that the hardware and software is set up as expected.”

— NVIDIA DeepOps: Slurm Deployment Guide

Kubernetes GPU Ops: simulator lab

Objectives 3.2

Check the GPU Operator, request GPUs in Kubernetes, debug a Pending pod, and spot time-slicing oversubscription.

Open the lab simulator Pick “Kubernetes GPU Ops” from the lab list.

Try this

  • Request nvidia.com/gpu in a pod spec.
  • Read why a pod is Pending.
  • Run the CUDA sample workload after maintenance.
What NVIDIA says (7)

“With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”

— NVIDIA device plugin for Kubernetes

“Check that all GPU Operator pods are running:”

— GPU Operator: Installing the NVIDIA GPU Operator

“within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU.”

— About the NVIDIA GPU Operator

“nvidia.com/gpu: 1 # requesting 1 GPU”

— NVIDIA device plugin for Kubernetes

“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas, but for some workloads this is better than not being able to share at all.”

— Time-Slicing GPUs in Kubernetes

“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”

— NVIDIA Xid Catalog

“Verification: Running Sample GPU Applications”

— GPU Operator: Installing the NVIDIA GPU Operator

AI, ML & DL Foundations: simulator lab

Objectives 1.3, 1.4, 1.5, 1.8

Sort examples into AI, machine learning and deep learning, and see why GPUs suit deep learning.

Open the lab simulator Pick “AI, ML & DL Foundations” from the lab list.

Try this

  • Classify five workloads.
  • Explain why a deep learning workload benefits from a GPU.
What NVIDIA says (9)

“Deep learning is a subset of machine learning, with the difference that DL algorithms can automatically learn representations from data such as images, video, or text, without introducing human domain knowledge.”

— What Is Deep Learning and Why Does It Matter?

“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”

— CUDA Programming Guide: Introduction

“In contrast, a GPU is composed of hundreds of cores that can handle thousands of threads simultaneously.”

— What's the Difference Between a CPU and a GPU?

“As a subset of AI, machine learning in its most elemental form uses algorithms to parse data, learn from it, and then make predictions or determinations about something in the real world.”

— What is Machine Learning and Why Does It Matter?

“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”

— CUDA Programming Guide: Introduction

“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”

— CUDA Programming Guide: Introduction

“Architecturally, the CPU is composed of just a few cores with lots of cache memory that can handle a few software threads at a time.”

— What's the Difference Between a CPU and a GPU?

“This parallelism maps naturally to GPUs , providing a significant computation speedup over CPU-only training”

— What Is Deep Learning and Why Does It Matter?

“Tensor Cores were introduced in the NVIDIA Volta™ GPU architecture to accelerate matrix multiply and accumulate operations for machine learning and scientific applications.”

— GPU Performance Background User's Guide

Training vs Inference Serving: simulator lab

Objectives 1.2, 1.7

Compare a training job with an inference service: what each needs and how latency and throughput are traded.

Open the lab simulator Pick “Training vs Inference Serving” from the lab list.

Try this

  • Profile a training step and an inference request.
  • Change batch size and watch latency and throughput.
What NVIDIA says (11)

“That’s why inference optimization techniques are so critical: they allow enterprises to optimize throughput and latency so they can meet their service level agreements across a variety of use cases.”

— What's the Difference Between Deep Learning Training and Inference?

“Over many iterations, it converges on the correct weights, enabling it to reliably suggest accurate, useful code completions or translations.”

— What's the Difference Between Deep Learning Training and Inference?

“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”

— What's the Difference Between Deep Learning Training and Inference?

“Triton Inference Server enables teams to deploy any AI model from multiple deep learning and machine learning frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL, and more.”

— NVIDIA Triton Inference Server

“Concurrent model execution Dynamic batching”

— NVIDIA Triton Inference Server

“is an SDK for optimizing deep learning inference on NVIDIA GPUs.”

— NVIDIA TensorRT Documentation

“optimizes inference using quantization, layer and tensor fusion, and kernel tuning techniques.”

— NVIDIA TensorRT (developer page)

“It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”

— NVIDIA TensorRT Documentation

“TensorRT supports mixed precision”

— NVIDIA TensorRT Documentation

“TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy”

— NVIDIA TensorRT (developer page)

“Inference can’t happen without training.”

— What's the Difference Between Deep Learning Training and Inference?

NVIDIA AI Software Stack: simulator lab

Objectives 1.1, 1.6, 1.7

Walk the NVIDIA AI software stack from driver and CUDA up to libraries, containers and inference services.

Open the lab simulator Pick “NVIDIA AI Software Stack” from the lab list.

Try this

  • List which layer each tool belongs to.
  • Match a workload to the NVIDIA solution that serves it.
What NVIDIA says (13)

“to enable any computational workload to use the throughput capability of GPUs independent of graphics APIs.”

— CUDA Programming Guide: Introduction

“Libraries like cuBLAS, cuFFT, cuDNN, and CUTLASS are just a few examples of libraries that help developers avoid reimplementing well-established algorithms.”

— CUDA Programming Guide: Introduction

“NVIDIA CUDA-X™, built on CUDA, is a collection of libraries that deliver dramatically higher performance across application domains, including AI and HPC.”

— CUDA Platform for Accelerated Computing

“The NVIDIA CUDA Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives for deep neural networks.”

— NVIDIA cuDNN Documentation

“NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”

— Overview of NCCL — NCCL 2.32.3 documentation

“NCCL implements both collective communication and point-to-point send/receive primitives.”

— Overview of NCCL — NCCL 2.32.3 documentation

“All of them also include the needed GPU libraries, configuration files, and tools to rebuild the container.”

— NVIDIA Deep Learning Frameworks User Guide (NGC containers)

“NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI”

— NVIDIA NeMo Framework

“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”

— NVIDIA NIM Documentation

“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”

— NVIDIA NIM Microservices

“is a GPU-accelerated library for tabular data processing.”

— NVIDIA cuDF Documentation

“throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”

— What Is MLOps? (NVIDIA Glossary)

“DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the”

— NVIDIA DALI User Guide

AI Infrastructure Planning: simulator lab

Objectives 2.1, 2.2, 2.3, 2.6

Size racks for DGX H100 systems: power per rack, rack density, and the extra racks needed when power or cooling is limited.

Open the lab simulator Pick “AI Infrastructure Planning” from the lab list.

Try this

  • Work out the power for one, two and four systems per rack.
  • Decide how to deploy one SU when cooling is oversubscribed.
What NVIDIA says (13)

“System power consumption 10.2 kW max”

— DGX SuperPOD Data Center Design (H100): Planning

“However, rack densities can be customized to fit within the available power and cooling capacities at the data center.”

— DGX SuperPOD Data Center Design (H100): Planning

“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”

— DGX SuperPOD H100 Reference Architecture: Abstract

“DGX H100 systems are optimally deployed at a rack density of four systems per rack.”

— DGX SuperPOD Data Center Design (H100): Planning

“4 8 326.4 kW 40.8 kW”

— DGX SuperPOD Data Center Design (H100): Planning

“Oversubscribed cooling can sometimes be mitigated by lowering the rack density”

— DGX SuperPOD Data Center Design (H100): Cooling

“2 16 326.4 kW 20.4 kW”

— DGX SuperPOD Data Center Design (H100): Planning

“In this example, the resource constraint in cooling is “paid for” using floor space.”

— DGX SuperPOD Data Center Design (H100): Planning

“In either case, the main benefit of aisle containment is the prevention of air recirculation from the hot aisle to the cold aisle”

— DGX SuperPOD Data Center Design (H100): Cooling

“Four of the six power supplies must be energized for the system to operate.”

— DGX SuperPOD Data Center Design (H100): Electrical

“Due to this requirement, the data center must minimally provide N+1 power, where N equals two power sources.”

— DGX SuperPOD Data Center Design (H100): Electrical

“Each power source must be sized to support 50% of the total peak load.”

— DGX SuperPOD Data Center Design (H100): Electrical

“The three main resource constraints in an air-cooled data center environment are power, cooling, and space.”

— DGX SuperPOD Data Center Design (H100): Planning

DPU Offload & Cloud vs On-Prem: simulator lab

Objectives 2.4, 2.10

Decide when to offload infrastructure work to a DPU, and compare cloud and on-prem costs over time.

Open the lab simulator Pick “DPU Offload & Cloud vs On-Prem” from the lab list.

Try this

  • Identify host CPU work a DPU can take over.
  • Compare up-front and ongoing costs for a multi-year workload.
What NVIDIA says (13)

“The NVIDIA DOCA™ Framework enables rapidly creating and managing applications and services on top of the BlueField networking platform, leveraging industry-standard APIs.”

— NVIDIA DOCA Overview

“by harnessing the power of NVIDIA's BlueField data-processing units (DPUs) and SuperNICs”

— NVIDIA DOCA Overview

“IT leaders should evaluate the total cost of ownership (TCO) over time and consider factors such as data storage, compute resources, and ongoing maintenance.”

— What Is AI Infrastructure? (NVIDIA Glossary)

“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”

— CUDA Programming Guide: Introduction

“It offloads demanding work that can bog down CPUs, processors that typically execute tasks in serial fashion.”

— What Is Accelerated Computing?

“The BlueField-3 platforms integrate x8 / x16 Armv8.2+ A78 Hercules cores (64-bit)”

— NVIDIA BlueField-3 Networking Platform User Guide: Introduction

“A DPU is a system on a chip, or SoC, that combines: An industry-standard, high-performance, software-programmable, multi-core CPU, typically based on the widely used Arm architecture, tightly coupled to the other SoC components.”

— What Is a DPU? (NVIDIA Blog)

“Data packet parsing, matching and manipulation to implement an open virtual switch (OVS)”

— What Is a DPU? (NVIDIA Blog)

“BlueField-3 DPUs offload, accelerate, and isolate software-defined networking, storage, security, and management functions, significantly enhancing data center performance, efficiency, and security.”

— NVIDIA BlueField-3 Networking Platform User Guide: Introduction

“By decoupling data center infrastructure from business applications, BlueField-3 creates a secure, zero-trust data center infrastructure”

— NVIDIA BlueField-3 Networking Platform User Guide: Introduction

“Cloud-based solutions offer a cost-effective way to start AI initiatives by reducing acquisition costs and shifting capital expenditures (CapEx) to operational expenditures (OpEx).”

— What Is AI Infrastructure? (NVIDIA Glossary)

“Yet, while cloud solutions may have lower initial costs, long-term expenses can add up.”

— What Is AI Infrastructure? (NVIDIA Glossary)

“Specialists in moving data in data centers, DPUs, or data processing units, are a new class of programmable processor and will join CPUs and GPUs as one of the three pillars of computing.”

— What Is a DPU? (NVIDIA Blog)

GPU Virtualization: simulator lab

Objectives 3.4

Share GPUs across virtual machines with NVIDIA vGPU and compare with MIG and time-slicing.

Open the lab simulator Pick “GPU Virtualization” from the lab list.

Try this

  • Pick a vGPU profile.
  • Compare isolation with time-slicing.
What NVIDIA says (6)

“NVIDIA virtual GPU software enables multiple virtual machines (VMs) to have simultaneous, direct access to a single physical GPU”

— NVIDIA Virtual GPU Software Documentation

“Time-slicing trades the memory and fault-isolation that is provided by MIG for the ability to share a GPU by a larger number of users.”

— Time-Slicing GPUs in Kubernetes

“Display information on GRID virtual GPUs.”

— nvidia-smi documentation

“With MIG, each instance’s processors have separate and isolated paths through the entire memory system”

— MIG User Guide: Introduction

“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas, but for some workloads this is better than not being able to share at all.”

— Time-Slicing GPUs in Kubernetes

“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”

— Time-Slicing GPUs in Kubernetes