1.5 Power and cooling
PSU redundancy, N+1 feeds, power capping and matching cooling to the full heat load.
Key points
A PSU (power supply unit) converts rack power for the server. The DGX H100 has six. Four must be energized for the system to operate. That is why NVIDIA requires at least N+1 power with two power sources, each sized for 50% of peak load.
What NVIDIA says (2)
“Four of the six power supplies must be energized for the system to operate.”
“the data center must minimally provide N+1 power, where N equals two power sources. Each power source must be sized to support 50% of the total peak load.”
Hot aisle / cold aisle means racks face each other so cool air enters the fronts and hot air leaves the backs into a separate aisle. Containment walls off one aisle. Blanking panels cover empty rack spaces so hot air cannot leak back. NVIDIA says to optimize airflow first, before more drastic cooling mitigations.
What NVIDIA says (3)
“Before considering more drastic cooling mitigations, it is important to make sure that the airflow in the space is optimized and well managed.”
“unoccupied RU spaces in the rack should be covered with blanking panels”
“Rack densities greater than 4 DGX H100 Systems are not recommended, due to thermodynamic considerations”
TGP (Total Graphics Power) is the GPU's power budget. SMBPBI is an out-of-band channel that the BMC can use to set GPU power. The GPU's PMU (Performance Monitoring Unit) applies the most conservative limit.
What NVIDIA says (2)
“The GPU has three sources of power limits:”
“The GPU Performance Monitoring Unit (PMU) selects the most conservative policy to cap power”
Each power circuit is sized for 50% of the load. With N+1, N is two circuits. So the real heat load equals two circuits. Planning cooling for one circuit leaves the rack short.
What NVIDIA says (2)
“However, with the specified N+1 power provisioning, N equals two circuits.”
“Therefore, it is critical to align the cooling capacity with N, and not simply the capacity of a single power circuit.”
Three-phase power carries more power per circuit than single-phase. N+1 adds a spare circuit so one feed can fail. NVIDIA prefers 415 VAC, 32 A, three-phase, N+1 for dense racks.
What NVIDIA says (1)
“The preferred power for high-density deployment patterns is 415 VAC, 32A, three-phase, N+1.”
Key terms
- Redfish: DMTF's standard REST API for managing and monitoring a server through its BMC.
- N+1 power: Power provisioning with one more feed than the load needs, so one feed can fail without stopping the system.
- Rack power distribution unit: The power strip in a rack that feeds each server, usually fed from three-phase power.
- Power capping: Setting a maximum power draw for a GPU or system; the GPU applies the most conservative limit it receives.
- Blanking panel: A cover for empty rack units so hot exhaust air cannot loop back to the server intakes.
Sample question
How many of a DGX H100's six power supplies must be energized for it to run, according to the SuperPOD data center design guide?
Show the answer
Answer: Four
A PSU (power supply unit) converts rack power for the server. The DGX H100 has six. Four must be energized for the system to operate. That is why NVIDIA requires at least N+1 power with two power sources, each sized for 50% of peak load.
What NVIDIA says (2)
“Four of the six power supplies must be energized for the system to operate.”
“the data center must minimally provide N+1 power, where N equals two power sources. Each power source must be sized to support 50% of the total peak load.”
Practice 1.5 (5 questions) Full System and Server Bring-up guide
← 1.4 Firmware upgrades and fault detection · 1.6 Installing GPU servers →