1.2 Network topologies for AI factories

NCP-AII · System and Server Bring-up (31% of the exam) · Official objective: “Describe network topologies for AI factories.”

The four SuperPOD fabrics, scalable units and the rail-optimized compute fabric.

Key points

  1. A rail-optimized fabric connects GPU number N in every node to the same leaf switch, called a rail. A scalable unit (SU) is a building block of 32 nodes. Within an SU, same-rail traffic is one hop away. Traffic between rails goes through the spine.

    What NVIDIA says (3)

    “Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “Traffic between nodes, or between rails, traverses the spine layer.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems”

    — DGX SuperPOD H100 Reference Architecture: Abstract

  2. A fabric is a network of switches, cables and adapters. SuperPOD keeps traffic types apart. Compute carries GPU-to-GPU traffic. Storage carries file data. In-band management carries cluster services. Out-of-band (OOB) management reaches the BMCs and device management ports.

    What NVIDIA says (2)

    “DGX SuperPOD configurations utilize four network fabrics:”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “-Compute Fabric | -Storage Fabric | -In-Band Management Network | -Out-of-Band Management Network”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

  3. In-band means the normal OS network of each node, as opposed to the separate BMC network. In SuperPOD it carries cluster services such as BCM and Slurm, and reaches outside services such as the NGC registry.

    What NVIDIA says (2)

    “Provides connectivity for the in-cluster services such as Base Command Manager, Slurm and to other services outside of the cluster such as the NGC registry, code repositories, and data sources.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “Enables access to the home filesystem and storage pool.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

  4. OOB ports include BMCs, PDUs and switch management ports. A PDU (power distribution unit) feeds power to a rack. Nobody runs jobs over these ports, and they control the hardware. So they sit on a separate, secured network.

    What NVIDIA says (2)

    “It connects the management ports of all devices including DGX and management servers, storage, networking gear, rack PDUs, and all other devices.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

    “These are separate onto their own fabric because there is no use-case where users need access to these ports and are secured using logical network separation.”

    — DGX SuperPOD H100 Reference Architecture: Network Fabrics

Key terms

Sample question

In the DGX SuperPOD H100 reference architecture, how does traffic flow between two DGX nodes in the same scalable unit on the same rail?

Show the answer

Answer: It is one hop away, through the rail's leaf switch.

A rail-optimized fabric connects GPU number N in every node to the same leaf switch, called a rail. A scalable unit (SU) is a building block of 32 nodes. Within an SU, same-rail traffic is one hop away. Traffic between rails goes through the spine.

What NVIDIA says (3)

“Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

“Traffic between nodes, or between rails, traverses the spine layer.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems”

— DGX SuperPOD H100 Reference Architecture: Abstract

Practice 1.2 (4 questions) Full System and Server Bring-up guide

← 1.1 Deployment and validation sequence · 1.3 BMC, OOB and TPM setup →