2.3 Power and cooling

NCA-AIIO · AI Infrastructure (40% of the exam) · Official objective: “Identify key concepts, and high-level specifications related to power and cooling requirements within a datacenter”

How much power DGX systems draw and how cooling must match it.

Key points

  1. NVIDIA's table lists 40.8 kW of total power per rack for four DGX H100 systems. In an air-cooled data center, the main resource constraints are power, cooling and space. If cooling capacity is limited, rack power density is limited too, and the systems must spread across more racks.

    What NVIDIA says (3)

    “4 8 326.4 kW 40.8 kW”

    — DGX SuperPOD Data Center Design (H100): Planning

    “The three main resource constraints in an air-cooled data center environment are power, cooling, and space.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “For example, limitations in cooling capacity could constrain rack power density, resulting in a need to occupy a larger number of racks to house a given number of servers.”

    — DGX SuperPOD Data Center Design (H100): Planning

  2. Power draw drives two things: the electrical circuits feeding the rack and the cooling needed to remove the heat. The data center design guide lists DGX H100 system power consumption as 10.2 kW max. Four systems in one rack therefore need about 40.8 kW.

    What NVIDIA says (2)

    “System power consumption 10.2 kW max”

    — DGX SuperPOD Data Center Design (H100): Planning

    “4 8 326.4 kW 40.8 kW”

    — DGX SuperPOD Data Center Design (H100): Planning

  3. Every watt a GPU uses becomes heat that must be removed. NVIDIA's AI infrastructure glossary says high-density compute with large power and cooling requirements needs mechanical, electrical and liquid cooling systems, plus management software. Note that the DGX H100 itself is listed as air-cooled, so liquid cooling is a density decision, not a rule for every GPU server.

    What NVIDIA says (2)

    “When using high density compute with large power and cooling requirements, it requires mechanical, electrical, and liquid cooling systems with management software to run it all efficiently.”

    — What Is AI Infrastructure? (NVIDIA Glossary)

    “Rack units 8 Cooling Air”

    — DGX SuperPOD Data Center Design (H100): Planning

  4. N+1 means you have the N power sources you need plus one spare. For DGX H100 racks, N equals two circuits, each sized for 50% of peak load. NVIDIA warns that cooling must be aligned with N and the full heat load, not simply one circuit's capacity.

    What NVIDIA says (2)

    “It is critical to plan for the full heat load of the rack profiles, keeping in mind that the power provisioning is based on circuits that provide only 50% of the full load.”

    — DGX SuperPOD Data Center Design (H100): Cooling

    “Therefore, it is critical to align the cooling capacity with N, and not simply the capacity of a single power circuit.”

    — DGX SuperPOD Data Center Design (H100): Cooling

Key terms

Try it

Sample question

Using NVIDIA's table, how much power does a rack of four DGX H100 systems need, and what usually limits density?

Show the answer

Answer: About 40.8 kW per rack. Power and cooling capacity usually limit how many systems you can place per rack.

NVIDIA's table lists 40.8 kW of total power per rack for four DGX H100 systems. In an air-cooled data center, the main resource constraints are power, cooling and space. If cooling capacity is limited, rack power density is limited too, and the systems must spread across more racks.

What NVIDIA says (3)

“4 8 326.4 kW 40.8 kW”

— DGX SuperPOD Data Center Design (H100): Planning

“The three main resource constraints in an air-cooled data center environment are power, cooling, and space.”

— DGX SuperPOD Data Center Design (H100): Planning

“For example, limitations in cooling capacity could constrain rack power density, resulting in a need to occupy a larger number of racks to house a given number of servers.”

— DGX SuperPOD Data Center Design (H100): Planning

Practice 2.3 (4 questions) Full AI Infrastructure guide

← 2.2 Scaling GPU infrastructure · 2.4 On-prem vs cloud →