2.8 Data center networking protocols
RDMA, GPUDirect RDMA, InfiniBand subnet management and oversubscription.
Key points
RDMA (remote direct memory access) lets a network adapter move data straight into another machine's memory without the CPU copying it. NVIDIA describes InfiniBand as RDMA-capable. In DGX SuperPOD, storage uses RDMA over InfiniBand to maximize performance and minimize CPU overhead.
What NVIDIA says (2)
“InfiniBand is a high-performance, low latency, RDMA capable networking technology”
“Storage is provided over InfiniBand and leverages RDMA to provide maximum performance and minimize CPU overhead.”
Ordinary RDMA moves data into host (CPU) memory. GPUDirect RDMA lets a peer device, such as a network adapter, read and write GPU memory directly using standard PCI Express features. NVIDIA says this provides direct communication between GPUs in remote systems.
What NVIDIA says (3)
“GPUDirect RDMA is a technology introduced in Kepler-class GPUs and CUDA 5.0 that enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”
“enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”
“Designed specifically for the needs of GPU acceleration, GPUDirect RDMA provides direct communication between NVIDIA GPUs in remote systems.”
A Subnet Manager (SM) configures the InfiniBand fabric. It discovers devices and sets up routing. NVIDIA's OFED documentation says all InfiniBand-compliant upper-layer protocols need a Subnet Manager running on the fabric at all times. OpenSM is one, installed with NVIDIA OFED. NVIDIA UFM can also manage the fabric.
What NVIDIA says (2)
“InfiniBand Subnet Manager All InfiniBand-compliant ULPs require a proper operation of a Subnet Manager (SM) running on the InfiniBand fabric, at all times.”
“OpenSM is an InfiniBand-compliant Subnet Manager, and it is installed as part of NVIDIA OFED”
Two things are needed. First, a network adapter that can do RDMA. Second, GPUDirect RDMA so that adapter can reach GPU memory directly. NVIDIA says network adapters can then read and write GPU memory directly, removing extra copies and reducing CPU overhead and latency.
What NVIDIA says (3)
“enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”
“Examples of third-party devices are: network interfaces, video acquisition devices, storage adapters.”
“avoids a bounce buffer through the CPU. Using this direct path can relieve system bandwidth bottlenecks and decrease the latency and utilization load on the CPU.”
Oversubscription means more bandwidth can enter a switch layer than can leave it. A ratio near 4:3 means about four units of possible demand for every three units of capacity. NVIDIA chose this for the storage fabric to allow more flexibility on cost and performance. The compute fabric, by contrast, is a full fat-tree.
What NVIDIA says (2)
“The DGX H100 system connections are slightly oversubscribed with a ratio near 4:3 with adjustments as needed to enable more storage flexibility regarding cost and performance.”
“Rail-optimized, full fat-tree network with eight NDR400 connections per system”
Key terms
- Remote direct memory access: Moving data straight into another machine's memory without the CPU copying it, which cuts latency and CPU load.
- GPUDirect RDMA: A direct path between GPU memory and a peer device such as a network adapter over PCI Express.
- RDMA over Converged Ethernet: RDMA carried over Ethernet networks instead of InfiniBand.
- InfiniBand: A high-performance, low-latency, RDMA-capable network used to connect GPU servers and storage.
- Subnet Manager: The service that configures an InfiniBand fabric; one must be running at all times.
- Oversubscription: When more bandwidth (or demand) can enter a layer than its capacity can serve.
- Fabric: The network that connects the servers: the switches, cables and adapters together.
- Lossless Ethernet: Ethernet configured to pause traffic instead of dropping packets when buffers fill, which RoCE needs.
Try it
Sample question
What is RDMA, and why do AI fabrics rely on it?
Show the answer
Answer: Remote direct memory access. It lets devices read and write memory directly, cutting copies, CPU overhead and latency.
RDMA (remote direct memory access) lets a network adapter move data straight into another machine's memory without the CPU copying it. NVIDIA describes InfiniBand as RDMA-capable. In DGX SuperPOD, storage uses RDMA over InfiniBand to maximize performance and minimize CPU overhead.
What NVIDIA says (2)
“InfiniBand is a high-performance, low latency, RDMA capable networking technology”
“Storage is provided over InfiniBand and leverages RDMA to provide maximum performance and minimize CPU overhead.”
Practice 2.8 (5 questions) Full AI Infrastructure guide
← 2.7 Networking requirements for AI · 2.9 High-speed network options →