2.7 Networking requirements for AI
Separate fabrics, rail-optimized design and the NCCL settings that pick a network path.
Key points
NCCL is the NVIDIA Collective Communications Library. Frameworks use it to exchange gradients between GPUs. NCCL_IB_DISABLE prevents NCCL from using the IB/RoCE transport. NCCL then falls back to another transport such as IP sockets, which is much slower. Unset it, or check that nothing in the job environment sets it.
What NVIDIA says (2)
“The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”
“NCCL will instead fall back to another available transport such as IP sockets.”
An HCA (host channel adapter) is an InfiniBand network adapter. NCCL_IB_HCA tells NCCL which RDMA interfaces to use. NCCL_SOCKET_IFNAME does the same job for IP interfaces.
What NVIDIA says (2)
“The NCCL_IB_HCA variable specifies which Host Channel Adapter (RDMA) interfaces to use for communication.”
“The NCCL_SOCKET_IFNAME variable specifies which IP interfaces to use for communication.”
NVLink is NVIDIA's direct GPU-to-GPU interconnect inside a server. Between servers, DGX SuperPOD uses NDR InfiniBand. NCCL, NVIDIA's library for multi-GPU communication, supports NVLink, PCIe, InfiniBand and IP sockets, and picks paths within and across nodes.
What NVIDIA says (3)
“is a direct GPU-to-GPU interconnect that scales multi-GPU input/output (IO) in the server.”
“NVIDIA NDR (400 Gbps) InfiniBand—bringing the highest performance, lowest latency, and most scalable network interconnect.”
“It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”
A rail is the set of same-numbered network ports across all nodes. Port 1 of every node goes to one leaf switch, port 2 to another, and so on. NVIDIA says the compute fabric is rail-optimized and each group of 32 nodes is rail-aligned. Traffic per rail is always one hop away from the other 31 nodes in an SU. Traffic between rails uses the spine layer.
What NVIDIA says (3)
“The compute fabric is rail-optimized to the top layer of the fabric.”
“Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”
“Traffic between nodes, or between rails, traverses the spine layer.”
NDR is the 400 Gb/s generation of InfiniBand. The reference architecture specifies a rail-optimized, full fat-tree network with eight NDR400 connections per system. The DGX H100 user guide shows eight ConnectX-7 single-port InfiniBand cards. With eight GPUs, that gives each GPU its own path onto its rail.
What NVIDIA says (2)
“Rail-optimized, full fat-tree network with eight NDR400 connections per system”
“Network (Cluster) card 4 x OSFP ports for 8 x NVIDIA® ConnectX®-7 Single Port InfiniBand Cards”
A fabric is a network built for one job. DGX SuperPOD uses four: the compute fabric for GPU-to-GPU traffic between nodes, the storage fabric, the in-band management network, and the out-of-band (OOB) management network. The OOB network connects the BMC (baseboard management controller) ports and stays physically isolated from users.
What NVIDIA says (2)
“DGX SuperPOD configurations utilize four network fabrics: -Compute Fabric -Storage Fabric -In-Band Management Network -Out-of-Band Management Network”
“The OOB management network connects all the base management controller (BMC) ports, as well as other devices that should be physically isolated from system users.”
Key terms
- NCCL: The library that moves data between GPUs inside and across nodes, with operations such as all-reduce.
- Rail-optimized network: A compute fabric where each same-numbered port on every node connects to its own set of leaf switches, keeping that traffic one hop apart.
Try it
Sample question
On an InfiniBand cluster, NCCL's log shows it is using IP sockets instead of InfiniBand, and training is slow. Which environment variable, if set, would cause exactly this?
Show the answer
Answer: NCCL_IB_DISABLE=1
NCCL is the NVIDIA Collective Communications Library. Frameworks use it to exchange gradients between GPUs. NCCL_IB_DISABLE prevents NCCL from using the IB/RoCE transport. NCCL then falls back to another transport such as IP sockets, which is much slower. Unset it, or check that nothing in the job environment sets it.
What NVIDIA says (2)
“The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”
“NCCL will instead fall back to another available transport such as IP sockets.”
Practice 2.7 (6 questions) Full AI Infrastructure guide
← 2.6 Facility requirements · 2.8 Data center networking protocols →