2.2 Scaling GPU infrastructure
How clusters grow in repeatable scalable units and what limits rack density.
Key points
A scalable unit (SU) is a pre-designed block of compute, networking and cabling that you repeat to grow a cluster. In the DGX SuperPOD H100 reference architecture, each SU contains 32 DGX H100 systems. Because every SU is the same, NVIDIA says deployment times drop from months to weeks.
What NVIDIA says (2)
“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”
“Using SUs, system deployment times are reduced from months to weeks.”
The reference architecture describes 4 SUs with 128 DGX nodes. It also says the design can scale to much larger configurations, up to and beyond 64 SUs with 2000+ DGX H100 nodes. Larger sizes add switch layers. For example, its table adds core switches at 16 SUs.
What NVIDIA says (2)
“This reference architecture is focused on 4 SU units with 128 DGX nodes.”
“DGX SuperPOD can scale to much larger configurations up to and beyond 64 SU with 2000+ DGX H100 nodes.”
UFM (Unified Fabric Manager) is NVIDIA's software for managing an InfiniBand fabric. The data center design guide says a pair of UFM appliances displaces one DGX H100 system. That leaves a maximum of 127 DGX H100 systems per full DGX SuperPOD.
What NVIDIA says (2)
“NVIDIA Unified Fabric Manager (UFM) appliances displaces one DGX H100 system in the DGX SuperPOD deployment pattern, resulting in a maximum of 127 DGX H100 systems per full DGX SuperPOD.”
“UFM enables data center operators to efficiently monitor and operate the entire fabric, boost application performance and maximize fabric resource utilization.”
Rack density is how many systems you put in one rack. NVIDIA says four DGX H100 systems per rack is optimal, but density can be customized to fit the available power and cooling. Lowering density increases the number of racks you need. The guide's table shows one SU always needs 326.4 kW of server power, whether it uses 8, 16 or 32 racks.
What NVIDIA says (4)
“DGX H100 systems are optimally deployed at a rack density of four systems per rack.”
“However, rack densities can be customized to fit within the available power and cooling capacities at the data center.”
“Naturally, reducing rack density increases the total number of required racks.”
“1 32 326.4 kW 10.2 kW 2 16 326.4 kW 20.4 kW 4 8 326.4 kW 40.8 kW”
Key terms
- DGX SuperPOD: NVIDIA's reference design for an AI cluster made of DGX systems, InfiniBand and Ethernet networking, management nodes and storage.
- Scalable unit: A repeatable building block of 32 DGX H100 systems used to grow a DGX SuperPOD.
- Rack density: How many systems are placed in one rack; it is limited by the power and cooling available per rack.
Try it
Sample question
DGX SuperPOD is built from 'scalable units' (SUs). What is an SU, and why use it?
Show the answer
Answer: A building block of 32 DGX H100 systems. Repeating it allows rapid deployment of systems of different sizes.
A scalable unit (SU) is a pre-designed block of compute, networking and cabling that you repeat to grow a cluster. In the DGX SuperPOD H100 reference architecture, each SU contains 32 DGX H100 systems. Because every SU is the same, NVIDIA says deployment times drop from months to weeks.
What NVIDIA says (2)
“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”
“Using SUs, system deployment times are reduced from months to weeks.”
Practice 2.2 (4 questions) Full AI Infrastructure guide
← 2.1 Hardware for training workloads · 2.3 Power and cooling →