AI Infrastructure

40% of the NCA-AIIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Essential AI Knowledge · AI Infrastructure · AI Operations

2.1 Hardware for training workloads

Official objective: “Identify hardware requirements for specific AI training task use cases”

What a training node provides: GPUs, GPU memory, local NVMe and fast networking.

Key points

  1. A cache is fast, nearby storage that keeps a copy of data so it does not have to be fetched again over the network. The DGX SuperPOD reference architecture rates storage needs by workload. It puts NLP at the 'Good' level because NLP datasets generally fit within the local cache. Larger datasets, such as compressed images, push the requirement higher.

    What NVIDIA says (2)

    “Good Natural Language Processing (NLP) Datasets generally fit within local cache”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

    “Ideally, data is cached during the first read of the dataset, so data does not have to be retrieved across the network.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

  2. GPU memory is the memory on the GPU itself. Model weights and intermediate values must fit in it during training. The DGX H100 user guide lists 8 x NVIDIA H100 GPUs that provide 640 GB total GPU memory. That is 80 GB per GPU. If a model needs more than that, it is split across GPUs, which is why the GPU-to-GPU link matters.

    What NVIDIA says (2)

    “The DGX H100/H200 systems are built on eight NVIDIA H100 Tensor Core GPUs or eight NVIDIA H200 Tensor Core GPUs.”

    — DGX H100/H200 User Guide: Introduction

    “For H100: 8 x NVIDIA H100 GPUs that provide 640 GB total GPU memory”

    — DGX H100/H200 User Guide: Introduction

  3. NVMe (Non-Volatile Memory Express) is a fast protocol for solid-state drives. A DGX H100 has a RAID 0 array of NVMe U.2 drives labeled as the data cache. The reference architecture says this local NVMe storage can be used for caching or staging data. Staging means copying data close to the GPUs before training starts.

    What NVIDIA says (2)

    “In addition, the DGX H100 system provides local NVMe storage that can also be used for caching or staging data.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

    “Storage (Data Cache) 8 x 3.84 TB NVMe U.2 SED (ea) in RAID 0 array”

    — DGX H100/H200 User Guide: Introduction

Try it: AI Infrastructure Planning

Practice 2.1 (3 questions)

2.2 Scaling GPU infrastructure

Official objective: “Scale a GPU infrastructure for different use cases”

How clusters grow in repeatable scalable units and what limits rack density.

Key points

  1. A scalable unit (SU) is a pre-designed block of compute, networking and cabling that you repeat to grow a cluster. In the DGX SuperPOD H100 reference architecture, each SU contains 32 DGX H100 systems. Because every SU is the same, NVIDIA says deployment times drop from months to weeks.

    What NVIDIA says (2)

    “The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”

    — DGX SuperPOD H100 Reference Architecture: Abstract

    “Using SUs, system deployment times are reduced from months to weeks.”

    — DGX SuperPOD H100 Reference Architecture: Components

  2. The reference architecture describes 4 SUs with 128 DGX nodes. It also says the design can scale to much larger configurations, up to and beyond 64 SUs with 2000+ DGX H100 nodes. Larger sizes add switch layers. For example, its table adds core switches at 16 SUs.

    What NVIDIA says (2)

    “This reference architecture is focused on 4 SU units with 128 DGX nodes.”

    — DGX SuperPOD H100 Reference Architecture: Architecture

    “DGX SuperPOD can scale to much larger configurations up to and beyond 64 SU with 2000+ DGX H100 nodes.”

    — DGX SuperPOD H100 Reference Architecture: Architecture

  3. UFM (Unified Fabric Manager) is NVIDIA's software for managing an InfiniBand fabric. The data center design guide says a pair of UFM appliances displaces one DGX H100 system. That leaves a maximum of 127 DGX H100 systems per full DGX SuperPOD.

    What NVIDIA says (2)

    “NVIDIA Unified Fabric Manager (UFM) appliances displaces one DGX H100 system in the DGX SuperPOD deployment pattern, resulting in a maximum of 127 DGX H100 systems per full DGX SuperPOD.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “UFM enables data center operators to efficiently monitor and operate the entire fabric, boost application performance and maximize fabric resource utilization.”

    — NVIDIA UFM Enterprise User Manual: Overview

  4. Rack density is how many systems you put in one rack. NVIDIA says four DGX H100 systems per rack is optimal, but density can be customized to fit the available power and cooling. Lowering density increases the number of racks you need. The guide's table shows one SU always needs 326.4 kW of server power, whether it uses 8, 16 or 32 racks.

    What NVIDIA says (4)

    “DGX H100 systems are optimally deployed at a rack density of four systems per rack.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “However, rack densities can be customized to fit within the available power and cooling capacities at the data center.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “Naturally, reducing rack density increases the total number of required racks.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “1 32 326.4 kW 10.2 kW 2 16 326.4 kW 20.4 kW 4 8 326.4 kW 40.8 kW”

    — DGX SuperPOD Data Center Design (H100): Planning

Key terms: DGX SuperPOD Scalable unit Rack density

Try it: AI Infrastructure Planning

Practice 2.2 (4 questions)

2.3 Power and cooling

Official objective: “Identify key concepts, and high-level specifications related to power and cooling requirements within a datacenter”

How much power DGX systems draw and how cooling must match it.

Key points

  1. NVIDIA's table lists 40.8 kW of total power per rack for four DGX H100 systems. In an air-cooled data center, the main resource constraints are power, cooling and space. If cooling capacity is limited, rack power density is limited too, and the systems must spread across more racks.

    What NVIDIA says (3)

    “4 8 326.4 kW 40.8 kW”

    — DGX SuperPOD Data Center Design (H100): Planning

    “The three main resource constraints in an air-cooled data center environment are power, cooling, and space.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “For example, limitations in cooling capacity could constrain rack power density, resulting in a need to occupy a larger number of racks to house a given number of servers.”

    — DGX SuperPOD Data Center Design (H100): Planning

  2. Power draw drives two things: the electrical circuits feeding the rack and the cooling needed to remove the heat. The data center design guide lists DGX H100 system power consumption as 10.2 kW max. Four systems in one rack therefore need about 40.8 kW.

    What NVIDIA says (2)

    “System power consumption 10.2 kW max”

    — DGX SuperPOD Data Center Design (H100): Planning

    “4 8 326.4 kW 40.8 kW”

    — DGX SuperPOD Data Center Design (H100): Planning

  3. Every watt a GPU uses becomes heat that must be removed. NVIDIA's AI infrastructure glossary says high-density compute with large power and cooling requirements needs mechanical, electrical and liquid cooling systems, plus management software. Note that the DGX H100 itself is listed as air-cooled, so liquid cooling is a density decision, not a rule for every GPU server.

    What NVIDIA says (2)

    “When using high density compute with large power and cooling requirements, it requires mechanical, electrical, and liquid cooling systems with management software to run it all efficiently.”

    — What Is AI Infrastructure? (NVIDIA Glossary)

    “Rack units 8 Cooling Air”

    — DGX SuperPOD Data Center Design (H100): Planning

  4. N+1 means you have the N power sources you need plus one spare. For DGX H100 racks, N equals two circuits, each sized for 50% of peak load. NVIDIA warns that cooling must be aligned with N and the full heat load, not simply one circuit's capacity.

    What NVIDIA says (2)

    “It is critical to plan for the full heat load of the rack profiles, keeping in mind that the power provisioning is based on circuits that provide only 50% of the full load.”

    — DGX SuperPOD Data Center Design (H100): Cooling

    “Therefore, it is critical to align the cooling capacity with N, and not simply the capacity of a single power circuit.”

    — DGX SuperPOD Data Center Design (H100): Cooling

Key terms: Rack density Aisle containment Oversubscription DGX system

Try it: AI Infrastructure Planning

Practice 2.3 (4 questions)

2.4 On-prem vs cloud

Official objective: “Articulate the key advantages, challenges, and considerations related to on-prem vs cloud infrastructures”

How to weigh up-front and ongoing costs, ownership and managed services.

Key points

  1. TCO (total cost of ownership) is the full cost over the life of the system, including storage, compute and maintenance. ROI (return on investment) compares what you gain with what you spend. NVIDIA says IT leaders should evaluate TCO over time and treat ROI, not the initial TCO, as a key metric.

    What NVIDIA says (2)

    “IT leaders should evaluate the total cost of ownership (TCO) over time and consider factors such as data storage, compute resources, and ongoing maintenance.”

    — What Is AI Infrastructure? (NVIDIA Glossary)

    “In general, it’s important to consider return on investment (ROI) as a key metric, rather than the initial TCO.”

    — What Is AI Infrastructure? (NVIDIA Glossary)

  2. On-premises (on-prem) means the hardware sits in a data center you control. NVIDIA says DGX SuperPOD can be deployed on-premises, meaning the customer owns and manages the hardware as a traditional system. The trade-off is that you also own power, cooling, space and operations.

    What NVIDIA says (1)

    “DGX SuperPOD can be deployed on-premises, meaning the customer owns and manages the hardware as a traditional system.”

    — DGX SuperPOD H100 Reference Architecture: Components

  3. A fully managed service means the provider runs the infrastructure and you use it. NVIDIA describes DGX Cloud with Microsoft Azure, AWS, Google Cloud and OCI as a high-performance, fully managed AI training platform. An on-prem DGX SuperPOD, by contrast, is owned and managed by the customer.

    What NVIDIA says (2)

    “NVIDIA DGX Cloud with Microsoft Azure is a high-performance, fully managed AI training platform”

    — NVIDIA DGX Cloud

    “DGX Cloud is built on NVIDIA-accelerated infrastructure and runs across CSPs and NVIDIA Cloud Partners (NCPs).”

    — NVIDIA DGX Cloud

  4. CapEx (capital expenditure) is money spent up front to buy assets such as servers. OpEx (operational expenditure) is ongoing spending, such as a monthly cloud bill. NVIDIA says cloud reduces acquisition costs and shifts CapEx to OpEx. It also warns that long-term cloud expenses can add up.

    What NVIDIA says (2)

    “Cloud-based solutions offer a cost-effective way to start AI initiatives by reducing acquisition costs and shifting capital expenditures (CapEx) to operational expenditures (OpEx).”

    — What Is AI Infrastructure? (NVIDIA Glossary)

    “Yet, while cloud solutions may have lower initial costs, long-term expenses can add up.”

    — What Is AI Infrastructure? (NVIDIA Glossary)

Key terms: CapEx and OpEx Total cost of ownership

Try it: DPU Offload & Cloud vs On-Prem

Practice 2.4 (4 questions)

2.5 Components of an accelerated cluster

Official objective: “Identify key components and considerations of a cluster of an accelerated infrastructure”

Compute, networking, management and storage, and how training uses storage.

Key points

  1. I/O (input/output) here means reading and writing data to storage. Training passes over the same dataset many times. Each pass is an epoch. NVIDIA says the key I/O operation in DL training is re-read. That is why caching the data after the first read helps so much.

    What NVIDIA says (2)

    “The key I/O operation in DL training is re-read.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

    “It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

  2. DMA (direct memory access) lets a device move data without the CPU copying each byte. A bounce buffer is an extra copy in CPU memory on the way to the GPU. GDS enables a direct DMA path between GPU memory and storage that avoids that bounce buffer. NVIDIA says this can relieve bandwidth bottlenecks and lower latency and CPU load.

    What NVIDIA says (2)

    “GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which avoids a bounce buffer through the CPU.”

    — GPUDirect Storage Overview

    “Using this direct path can relieve system bandwidth bottlenecks and decrease the latency and utilization load on the CPU.”

    — GPUDirect Storage Overview

  3. A checkpoint is a saved copy of the model's state during training. If a node fails, the job restarts from the last checkpoint instead of from the beginning. NVIDIA says writing checkpoints is necessary for fault tolerance as models grow, and that write performance can also be important.

    What NVIDIA says (2)

    “As DL models grow in size and time-to-train, writing checkpoints is necessary for fault tolerance.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

    “Write performance can also be important.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

  4. An accelerated cluster is more than its GPU servers. NVIDIA defines the DGX SuperPOD architecture as a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage. Each part has to keep up with the others or the GPUs sit idle.

    What NVIDIA says (1)

    “The DGX SuperPOD architecture is a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage.”

    — DGX SuperPOD H100 Reference Architecture: Architecture

Key terms: DGX SuperPOD GPUDirect Storage Checkpoint NVLink PCI Express Bandwidth

Try it: NVLink Topology XID Fault Drill Storage Bottleneck GPUDirect Storage

Practice 2.5 (4 questions)

2.6 Facility requirements

Official objective: “Identify facility requirements”

Redundancy standards, power feeds, floor space and airflow for an AI data center.

Key points

  1. A well-planned deployment thinks about growth from day one. NVIDIA says to reserve floor space for the future state of the system, not just the initial state. It also warns that performance-based cable length limits prohibit placing racks or SUs too far apart.

    What NVIDIA says (2)

    “A well-planned deployment will consider future expansion and reserve sufficient floor space for the future state of the system, not just the initial state.”

    — DGX SuperPOD Data Center Design (H100): Infrastructure

    “Performance-based cable length limitations prohibit distributing the racks or scalable units too far from one another.”

    — DGX SuperPOD Data Center Design (H100): Infrastructure

  2. A resource is oversubscribed when demand exceeds capacity. NVIDIA says oversubscribed cooling can be mitigated by lowering rack density or by spacing racks apart. In that case, the cooling shortfall is 'paid for' with floor space. It also stresses first optimizing airflow, for example with aisle containment, which stops hot exhaust air from recirculating to server inlets.

    What NVIDIA says (3)

    “Oversubscribed cooling can sometimes be mitigated by lowering the rack density thereby reducing the cooling demand per rack footprint, or by spacing the racks further apart to aggregate the cooling capacity of more than one rack footprint to each populated rack.”

    — DGX SuperPOD Data Center Design (H100): Cooling

    “In this example, the resource constraint in cooling is “paid for” using floor space.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “In either case, the main benefit of aisle containment is the prevention of air recirculation from the hot aisle to the cold aisle, which artificially increases the supply air temperature at the inlet of the servers, significantly reducing its heat exchange potential.”

    — DGX SuperPOD Data Center Design (H100): Cooling

  3. A PSU (power supply unit) converts facility power for the server. NVIDIA says four of the six PSUs must be energized for the system to operate. With only two feeds, losing one would drop a system to three PSUs and shut it down. So NVIDIA requires at least N+1 power with N equal to two sources, each sized for 50% of the total peak load.

    What NVIDIA says (5)

    “The system includes six internal power supply units.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Four of the six power supplies must be energized for the system to operate.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Due to this requirement, the data center must minimally provide N+1 power, where N equals two power sources.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Each power source must be sized to support 50% of the total peak load.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Not acceptable for DGX H100 systems”

    — DGX SuperPOD Data Center Design (H100): Electrical

  4. Concurrent maintainability means you can service any power or cooling part without shutting down the IT load. NVIDIA says the data center should generally meet or exceed Uptime Institute Tier 3, or an equivalent standard. That includes concurrent maintainability and no single point of failure.

    What NVIDIA says (1)

    “Generally, the data center should meet or exceed Uptime Institute Tier 3 design standards, or alternatively the TIA942-B Rated 3 or EN50600 Availability Class 3 design standards, including concurrent maintainability and no single point of failure.”

    — DGX SuperPOD Data Center Design (H100): Electrical

Key terms: N+1 redundancy Aisle containment Gradient

Try it: AI Infrastructure Planning

Practice 2.6 (4 questions)

2.7 Networking requirements for AI

Official objective: “Determine networking requirements for AI workloads”

Separate fabrics, rail-optimized design and the NCCL settings that pick a network path.

Key points

  1. NCCL is the NVIDIA Collective Communications Library. Frameworks use it to exchange gradients between GPUs. NCCL_IB_DISABLE prevents NCCL from using the IB/RoCE transport. NCCL then falls back to another transport such as IP sockets, which is much slower. Unset it, or check that nothing in the job environment sets it.

    What NVIDIA says (2)

    “The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”

    — NCCL Environment Variables

    “NCCL will instead fall back to another available transport such as IP sockets.”

    — NCCL Environment Variables

  2. An HCA (host channel adapter) is an InfiniBand network adapter. NCCL_IB_HCA tells NCCL which RDMA interfaces to use. NCCL_SOCKET_IFNAME does the same job for IP interfaces.

    What NVIDIA says (2)

    “The NCCL_IB_HCA variable specifies which Host Channel Adapter (RDMA) interfaces to use for communication.”

    — NCCL Environment Variables

    “The NCCL_SOCKET_IFNAME variable specifies which IP interfaces to use for communication.”

    — NCCL Environment Variables

  3. NVLink is NVIDIA's direct GPU-to-GPU interconnect inside a server. Between servers, DGX SuperPOD uses NDR InfiniBand. NCCL, NVIDIA's library for multi-GPU communication, supports NVLink, PCIe, InfiniBand and IP sockets, and picks paths within and across nodes.

    What NVIDIA says (3)

    “is a direct GPU-to-GPU interconnect that scales multi-GPU input/output (IO) in the server.”

    — NVIDIA Fabric Manager User Guide

    “NVIDIA NDR (400 Gbps) InfiniBand—bringing the highest performance, lowest latency, and most scalable network interconnect.”

    — DGX SuperPOD H100 Reference Architecture: Abstract

    “It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”

    — Overview of NCCL — NCCL 2.32.3 documentation

  4. A rail is the set of same-numbered network ports across all nodes. Port 1 of every node goes to one leaf switch, port 2 to another, and so on. NVIDIA says the compute fabric is rail-optimized and each group of 32 nodes is rail-aligned. Traffic per rail is always one hop away from the other 31 nodes in an SU. Traffic between rails uses the spine layer.

    What NVIDIA says (3)

    “The compute fabric is rail-optimized to the top layer of the fabric.”

    — DGX SuperPOD H100 Reference Architecture: Components

    “Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “Traffic between nodes, or between rails, traverses the spine layer.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

  5. NDR is the 400 Gb/s generation of InfiniBand. The reference architecture specifies a rail-optimized, full fat-tree network with eight NDR400 connections per system. The DGX H100 user guide shows eight ConnectX-7 single-port InfiniBand cards. With eight GPUs, that gives each GPU its own path onto its rail.

    What NVIDIA says (2)

    “Rail-optimized, full fat-tree network with eight NDR400 connections per system”

    — DGX SuperPOD H100 Reference Architecture: Components

    “Network (Cluster) card 4 x OSFP ports for 8 x NVIDIA® ConnectX®-7 Single Port InfiniBand Cards”

    — DGX H100/H200 User Guide: Introduction

  6. A fabric is a network built for one job. DGX SuperPOD uses four: the compute fabric for GPU-to-GPU traffic between nodes, the storage fabric, the in-band management network, and the out-of-band (OOB) management network. The OOB network connects the BMC (baseboard management controller) ports and stays physically isolated from users.

    What NVIDIA says (2)

    “DGX SuperPOD configurations utilize four network fabrics: -Compute Fabric -Storage Fabric -In-Band Management Network -Out-of-Band Management Network”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “The OOB management network connects all the base management controller (BMC) ports, as well as other devices that should be physically isolated from system users.”

    — DGX SuperPOD H100 Reference Architecture: Components

Key terms: NCCL Rail-optimized network

Try it: AllReduce Deep Dive NCCL Fallback Drill

Practice 2.7 (6 questions)

2.8 Data center networking protocols

Official objective: “Identify and describe DC networking protocols and key concepts”

RDMA, GPUDirect RDMA, InfiniBand subnet management and oversubscription.

Key points

  1. RDMA (remote direct memory access) lets a network adapter move data straight into another machine's memory without the CPU copying it. NVIDIA describes InfiniBand as RDMA-capable. In DGX SuperPOD, storage uses RDMA over InfiniBand to maximize performance and minimize CPU overhead.

    What NVIDIA says (2)

    “InfiniBand is a high-performance, low latency, RDMA capable networking technology”

    — DGX SuperPOD H100 Reference Architecture: Components

    “Storage is provided over InfiniBand and leverages RDMA to provide maximum performance and minimize CPU overhead.”

    — DGX SuperPOD H100 Reference Architecture: Components

  2. Ordinary RDMA moves data into host (CPU) memory. GPUDirect RDMA lets a peer device, such as a network adapter, read and write GPU memory directly using standard PCI Express features. NVIDIA says this provides direct communication between GPUs in remote systems.

    What NVIDIA says (3)

    “GPUDirect RDMA is a technology introduced in Kepler-class GPUs and CUDA 5.0 that enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

    — GPUDirect RDMA

    “enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

    — GPUDirect RDMA

    “Designed specifically for the needs of GPU acceleration, GPUDirect RDMA provides direct communication between NVIDIA GPUs in remote systems.”

    — NVIDIA GPUDirect (developer page)

  3. A Subnet Manager (SM) configures the InfiniBand fabric. It discovers devices and sets up routing. NVIDIA's OFED documentation says all InfiniBand-compliant upper-layer protocols need a Subnet Manager running on the fabric at all times. OpenSM is one, installed with NVIDIA OFED. NVIDIA UFM can also manage the fabric.

    What NVIDIA says (2)

    “InfiniBand Subnet Manager All InfiniBand-compliant ULPs require a proper operation of a Subnet Manager (SM) running on the InfiniBand fabric, at all times.”

    — MLNX_OFED Documentation: Introduction

    “OpenSM is an InfiniBand-compliant Subnet Manager, and it is installed as part of NVIDIA OFED”

    — MLNX_OFED Documentation: Introduction

  4. Two things are needed. First, a network adapter that can do RDMA. Second, GPUDirect RDMA so that adapter can reach GPU memory directly. NVIDIA says network adapters can then read and write GPU memory directly, removing extra copies and reducing CPU overhead and latency.

    What NVIDIA says (3)

    “enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

    — GPUDirect RDMA

    “Examples of third-party devices are: network interfaces, video acquisition devices, storage adapters.”

    — GPUDirect RDMA

    “avoids a bounce buffer through the CPU. Using this direct path can relieve system bandwidth bottlenecks and decrease the latency and utilization load on the CPU.”

    — GPUDirect Storage Overview

  5. Oversubscription means more bandwidth can enter a switch layer than can leave it. A ratio near 4:3 means about four units of possible demand for every three units of capacity. NVIDIA chose this for the storage fabric to allow more flexibility on cost and performance. The compute fabric, by contrast, is a full fat-tree.

    What NVIDIA says (2)

    “The DGX H100 system connections are slightly oversubscribed with a ratio near 4:3 with adjustments as needed to enable more storage flexibility regarding cost and performance.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “Rail-optimized, full fat-tree network with eight NDR400 connections per system”

    — DGX SuperPOD H100 Reference Architecture: Components

Key terms: Remote direct memory access GPUDirect RDMA RDMA over Converged Ethernet InfiniBand Subnet Manager Oversubscription Fabric Lossless Ethernet

Try it: InfiniBand Fabric RoCEv2 + PFC/ECN

Practice 2.8 (5 questions)

2.9 High-speed network options

Official objective: “Identify high speed DC network options and their use cases”

NVLink and NVLink Switch inside a system; InfiniBand and Spectrum-X Ethernet between systems.

Key points

  1. NVLink is NVIDIA's direct GPU-to-GPU interconnect. It scales multi-GPU input and output within a server. NVIDIA lists 1.5x higher bandwidth per GPU at 900 GBps with fourth-generation NVLink on DGX H100. Newer generations are faster, so always check which generation a question means.

    What NVIDIA says (2)

    “is a direct GPU-to-GPU interconnect that scales multi-GPU input/output (IO) in the server.”

    — NVIDIA Fabric Manager User Guide

    “1.5X higher bandwidth per GPU @ 900 GBps with fourth generation of NVIDIA NVLink.”

    — DGX SuperPOD H100 Reference Architecture: Components

  2. Point-to-point links only connect pairs. A switch lets every GPU reach every other GPU. NVIDIA says NVLink Switch chips connect multiple NVLinks to provide all-to-all GPU communication at full NVLink speed. NVLink Switches also include SHARP engines for in-network reductions.

    What NVIDIA says (3)

    “which connects multiple NVLinks to provide all-to-all GPU communication at the total NVLink speed.”

    — NVIDIA Fabric Manager User Guide

    “The NVIDIA NVLink Switch chips connect multiple NVLinks to provide all-to-all GPU communication at full NVLink speed across the entire rack.”

    — NVIDIA NVLink and NVLink Switch

    “To enable high-speed, collective operations, each NVLink Switch has engines for NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol (SHARP)™ for in-network reductions and multicast acceleration.”

    — NVIDIA NVLink and NVLink Switch

  3. Multi-tenant means many customers share the same infrastructure. 'Noise' is one tenant's traffic slowing another. NVIDIA says Spectrum-X Ethernet is designed for cloud providers and large enterprises running multi-tenant AI at hyperscale. It uses telemetry-based congestion control for noise isolation, and SuperNICs provide RoCE (RDMA over Converged Ethernet) between GPU servers.

    What NVIDIA says (4)

    “NVIDIA Spectrum-X is an AI-optimized Ethernet networking platform that combines NVIDIA Spectrum switches with the BlueField-3 SuperNIC”

    — NVIDIA Network Operator: Spectrum-X Ethernet Networking Platform

    “NVIDIA Spectrum-X Ethernet is designed for cloud service providers and large enterprises running multi-tenant AI workloads at hyperscale.”

    — NVIDIA Spectrum-X Ethernet

    “Spectrum-X Ethernet uses advanced telemetry-based congestion control to provide noise isolation between tenants.”

    — NVIDIA Spectrum-X Ethernet

    “the BlueField-3 SuperNIC provides best-in-class remote direct-memory access over converged Ethernet (RoCE) network connectivity between GPU servers”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

  4. SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) lets switches do part of a collective operation, so less data crosses the network. Adaptive routing sends traffic around busy links. NVIDIA highlights SHARP, quality-of-service features including congestion control and adaptive routing, and self-healing networking for Quantum InfiniBand.

    What NVIDIA says (4)

    “technology improves the performance of MPI and Machine Learning collective operation, by offloading collective operations from CPUs and GPUs to the network”

    — NVIDIA SHARP Documentation: Introduction

    “This innovative approach decreases the amount of data traversing the network as aggregation nodes are reached”

    — NVIDIA SHARP Documentation: Introduction

    “NVIDIA Quantum InfiniBand is the only high-performance interconnect solution with proven quality-of-service capabilities, including advanced congestion control and adaptive routing, resulting in unmatched network efficiency.”

    — NVIDIA InfiniBand Networking

    “NVIDIA Quantum InfiniBand with self-healing network capabilities overcomes link failures”

    — NVIDIA InfiniBand Networking

Key terms: RDMA over Converged Ethernet InfiniBand SHARP NVLink NVLink Switch Spectrum-X Ethernet

Try it: NVLink Topology XID Fault Drill InfiniBand Fabric

Practice 2.9 (4 questions)

2.10 DPUs in the data center

Official objective: “Explain the purpose and benefits of a DPU in a datacenter”

What a DPU is and how BlueField and DOCA offload, accelerate and isolate infrastructure work.

Key points

  1. Offload means moving work from one processor to another. NVIDIA says DOCA and BlueField let developers build services that offload, accelerate and isolate data center workloads. BlueField decouples data center infrastructure from business applications. Its acceleration engines cover functions such as open virtual switch (OVS) packet processing.

    What NVIDIA says (5)

    “The NVIDIA DOCA™ Framework enables rapidly creating and managing applications and services on top of the BlueField networking platform, leveraging industry-standard APIs.”

    — NVIDIA DOCA Overview

    “by harnessing the power of NVIDIA's BlueField data-processing units (DPUs) and SuperNICs”

    — NVIDIA DOCA Overview

    “BlueField-3 DPUs offload, accelerate, and isolate software-defined networking, storage, security, and management functions, significantly enhancing data center performance, efficiency, and security.”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

    “By decoupling data center infrastructure from business applications, BlueField-3 creates a secure, zero-trust data center infrastructure”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

    “Data packet parsing, matching and manipulation to implement an open virtual switch (OVS)”

    — What Is a DPU? (NVIDIA Blog)

  2. A DPU (data processing unit) is a processor built to move and process data in the data center. NVIDIA describes it as a system on a chip (SoC) with an Arm-based multi-core CPU, a high-performance network interface, and programmable acceleration engines. NVIDIA calls the DPU the third pillar of computing, next to CPUs and GPUs. NVIDIA's DPU family is BlueField.

    What NVIDIA says (4)

    “The BlueField-3 platforms integrate x8 / x16 Armv8.2+ A78 Hercules cores (64-bit)”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

    “A DPU is a system on a chip, or SoC, that combines: An industry-standard, high-performance, software-programmable, multi-core CPU, typically based on the widely used Arm architecture, tightly coupled to the other SoC components.”

    — What Is a DPU? (NVIDIA Blog)

    “BlueField-3 DPUs offload, accelerate, and isolate software-defined networking, storage, security, and management functions, significantly enhancing data center performance, efficiency, and security.”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

    “Specialists in moving data in data centers, DPUs, or data processing units, are a new class of programmable processor and will join CPUs and GPUs as one of the three pillars of computing.”

    — What Is a DPU? (NVIDIA Blog)

  3. An SDK (software development kit) is a set of libraries and tools for building software. BlueField is programmed through NVIDIA DOCA. DOCA contains a runtime and a development environment (the SDK). DOCA services are containerized DOCA programs that can be deployed from NVIDIA's NGC catalog directly to BlueField.

    What NVIDIA says (3)

    “all fully programmable through the NVIDIA DOCA™ software framework”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

    “DOCA contains a runtime and development environment, including libraries and drivers for device management and programmability”

    — NVIDIA DOCA Overview

    “DOCA services are accessible as part of NVIDIA's container catalog (NGC) from which they can be easily deployed directly to BlueField”

    — NVIDIA DOCA Overview

  4. A data path is the route network packets take through a system. NVIDIA lists capabilities a DPU's data-path accelerators need, including RDMA transport acceleration for Zero Touch RoCE and GPUDirect accelerators that bypass the CPU and feed networked data directly to GPUs.

    What NVIDIA says (2)

    “the BlueField-3 SuperNIC provides best-in-class remote direct-memory access over converged Ethernet (RoCE) network connectivity between GPU servers”

    — NVIDIA BlueField-3 Networking Platform User Guide: Introduction

    “RDMA data transport acceleration for Zero Touch RoCE GPUDirect accelerators to bypass the CPU and feed networked data directly to GPUs”

    — What Is a DPU? (NVIDIA Blog)

Key terms: Data processing unit DOCA

Try it: DPU Offload & Cloud vs On-Prem

Practice 2.10 (4 questions)