2.8 Data center networking protocols

NCA-AIIO · AI Infrastructure (40% of the exam) · Official objective: “Identify and describe DC networking protocols and key concepts”

RDMA, GPUDirect RDMA, InfiniBand subnet management and oversubscription.

Key points

  1. RDMA (remote direct memory access) lets a network adapter move data straight into another machine's memory without the CPU copying it. NVIDIA describes InfiniBand as RDMA-capable. In DGX SuperPOD, storage uses RDMA over InfiniBand to maximize performance and minimize CPU overhead.

    What NVIDIA says (2)

    “InfiniBand is a high-performance, low latency, RDMA capable networking technology”

    — DGX SuperPOD H100 Reference Architecture: Components

    “Storage is provided over InfiniBand and leverages RDMA to provide maximum performance and minimize CPU overhead.”

    — DGX SuperPOD H100 Reference Architecture: Components

  2. Ordinary RDMA moves data into host (CPU) memory. GPUDirect RDMA lets a peer device, such as a network adapter, read and write GPU memory directly using standard PCI Express features. NVIDIA says this provides direct communication between GPUs in remote systems.

    What NVIDIA says (3)

    “GPUDirect RDMA is a technology introduced in Kepler-class GPUs and CUDA 5.0 that enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

    — GPUDirect RDMA

    “enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

    — GPUDirect RDMA

    “Designed specifically for the needs of GPU acceleration, GPUDirect RDMA provides direct communication between NVIDIA GPUs in remote systems.”

    — NVIDIA GPUDirect (developer page)

  3. A Subnet Manager (SM) configures the InfiniBand fabric. It discovers devices and sets up routing. NVIDIA's OFED documentation says all InfiniBand-compliant upper-layer protocols need a Subnet Manager running on the fabric at all times. OpenSM is one, installed with NVIDIA OFED. NVIDIA UFM can also manage the fabric.

    What NVIDIA says (2)

    “InfiniBand Subnet Manager All InfiniBand-compliant ULPs require a proper operation of a Subnet Manager (SM) running on the InfiniBand fabric, at all times.”

    — MLNX_OFED Documentation: Introduction

    “OpenSM is an InfiniBand-compliant Subnet Manager, and it is installed as part of NVIDIA OFED”

    — MLNX_OFED Documentation: Introduction

  4. Two things are needed. First, a network adapter that can do RDMA. Second, GPUDirect RDMA so that adapter can reach GPU memory directly. NVIDIA says network adapters can then read and write GPU memory directly, removing extra copies and reducing CPU overhead and latency.

    What NVIDIA says (3)

    “enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

    — GPUDirect RDMA

    “Examples of third-party devices are: network interfaces, video acquisition devices, storage adapters.”

    — GPUDirect RDMA

    “avoids a bounce buffer through the CPU. Using this direct path can relieve system bandwidth bottlenecks and decrease the latency and utilization load on the CPU.”

    — GPUDirect Storage Overview

  5. Oversubscription means more bandwidth can enter a switch layer than can leave it. A ratio near 4:3 means about four units of possible demand for every three units of capacity. NVIDIA chose this for the storage fabric to allow more flexibility on cost and performance. The compute fabric, by contrast, is a full fat-tree.

    What NVIDIA says (2)

    “The DGX H100 system connections are slightly oversubscribed with a ratio near 4:3 with adjustments as needed to enable more storage flexibility regarding cost and performance.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “Rail-optimized, full fat-tree network with eight NDR400 connections per system”

    — DGX SuperPOD H100 Reference Architecture: Components

Key terms

Try it

Sample question

What is RDMA, and why do AI fabrics rely on it?

Show the answer

Answer: Remote direct memory access. It lets devices read and write memory directly, cutting copies, CPU overhead and latency.

RDMA (remote direct memory access) lets a network adapter move data straight into another machine's memory without the CPU copying it. NVIDIA describes InfiniBand as RDMA-capable. In DGX SuperPOD, storage uses RDMA over InfiniBand to maximize performance and minimize CPU overhead.

What NVIDIA says (2)

“InfiniBand is a high-performance, low latency, RDMA capable networking technology”

— DGX SuperPOD H100 Reference Architecture: Components

“Storage is provided over InfiniBand and leverages RDMA to provide maximum performance and minimize CPU overhead.”

— DGX SuperPOD H100 Reference Architecture: Components

Practice 2.8 (5 questions) Full AI Infrastructure guide

← 2.7 Networking requirements for AI · 2.9 High-speed network options →