2.2 Scaling GPU infrastructure

NCA-AIIO · AI Infrastructure (40% of the exam) · Official objective: “Scale a GPU infrastructure for different use cases”

How clusters grow in repeatable scalable units and what limits rack density.

Key points

  1. A scalable unit (SU) is a pre-designed block of compute, networking and cabling that you repeat to grow a cluster. In the DGX SuperPOD H100 reference architecture, each SU contains 32 DGX H100 systems. Because every SU is the same, NVIDIA says deployment times drop from months to weeks.

    What NVIDIA says (2)

    “The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”

    — DGX SuperPOD H100 Reference Architecture: Abstract

    “Using SUs, system deployment times are reduced from months to weeks.”

    — DGX SuperPOD H100 Reference Architecture: Components

  2. The reference architecture describes 4 SUs with 128 DGX nodes. It also says the design can scale to much larger configurations, up to and beyond 64 SUs with 2000+ DGX H100 nodes. Larger sizes add switch layers. For example, its table adds core switches at 16 SUs.

    What NVIDIA says (2)

    “This reference architecture is focused on 4 SU units with 128 DGX nodes.”

    — DGX SuperPOD H100 Reference Architecture: Architecture

    “DGX SuperPOD can scale to much larger configurations up to and beyond 64 SU with 2000+ DGX H100 nodes.”

    — DGX SuperPOD H100 Reference Architecture: Architecture

  3. UFM (Unified Fabric Manager) is NVIDIA's software for managing an InfiniBand fabric. The data center design guide says a pair of UFM appliances displaces one DGX H100 system. That leaves a maximum of 127 DGX H100 systems per full DGX SuperPOD.

    What NVIDIA says (2)

    “NVIDIA Unified Fabric Manager (UFM) appliances displaces one DGX H100 system in the DGX SuperPOD deployment pattern, resulting in a maximum of 127 DGX H100 systems per full DGX SuperPOD.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “UFM enables data center operators to efficiently monitor and operate the entire fabric, boost application performance and maximize fabric resource utilization.”

    — NVIDIA UFM Enterprise User Manual: Overview

  4. Rack density is how many systems you put in one rack. NVIDIA says four DGX H100 systems per rack is optimal, but density can be customized to fit the available power and cooling. Lowering density increases the number of racks you need. The guide's table shows one SU always needs 326.4 kW of server power, whether it uses 8, 16 or 32 racks.

    What NVIDIA says (4)

    “DGX H100 systems are optimally deployed at a rack density of four systems per rack.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “However, rack densities can be customized to fit within the available power and cooling capacities at the data center.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “Naturally, reducing rack density increases the total number of required racks.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “1 32 326.4 kW 10.2 kW 2 16 326.4 kW 20.4 kW 4 8 326.4 kW 40.8 kW”

    — DGX SuperPOD Data Center Design (H100): Planning

Key terms

Try it

Sample question

DGX SuperPOD is built from 'scalable units' (SUs). What is an SU, and why use it?

Show the answer

Answer: A building block of 32 DGX H100 systems. Repeating it allows rapid deployment of systems of different sizes.

A scalable unit (SU) is a pre-designed block of compute, networking and cabling that you repeat to grow a cluster. In the DGX SuperPOD H100 reference architecture, each SU contains 32 DGX H100 systems. Because every SU is the same, NVIDIA says deployment times drop from months to weeks.

What NVIDIA says (2)

“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”

— DGX SuperPOD H100 Reference Architecture: Abstract

“Using SUs, system deployment times are reduced from months to weeks.”

— DGX SuperPOD H100 Reference Architecture: Components

Practice 2.2 (4 questions) Full AI Infrastructure guide

← 2.1 Hardware for training workloads · 2.3 Power and cooling →