2.6 Facility requirements
Redundancy standards, power feeds, floor space and airflow for an AI data center.
Key points
A well-planned deployment thinks about growth from day one. NVIDIA says to reserve floor space for the future state of the system, not just the initial state. It also warns that performance-based cable length limits prohibit placing racks or SUs too far apart.
What NVIDIA says (2)
“A well-planned deployment will consider future expansion and reserve sufficient floor space for the future state of the system, not just the initial state.”
“Performance-based cable length limitations prohibit distributing the racks or scalable units too far from one another.”
A resource is oversubscribed when demand exceeds capacity. NVIDIA says oversubscribed cooling can be mitigated by lowering rack density or by spacing racks apart. In that case, the cooling shortfall is 'paid for' with floor space. It also stresses first optimizing airflow, for example with aisle containment, which stops hot exhaust air from recirculating to server inlets.
What NVIDIA says (3)
“Oversubscribed cooling can sometimes be mitigated by lowering the rack density thereby reducing the cooling demand per rack footprint, or by spacing the racks further apart to aggregate the cooling capacity of more than one rack footprint to each populated rack.”
“In this example, the resource constraint in cooling is “paid for” using floor space.”
“In either case, the main benefit of aisle containment is the prevention of air recirculation from the hot aisle to the cold aisle, which artificially increases the supply air temperature at the inlet of the servers, significantly reducing its heat exchange potential.”
A PSU (power supply unit) converts facility power for the server. NVIDIA says four of the six PSUs must be energized for the system to operate. With only two feeds, losing one would drop a system to three PSUs and shut it down. So NVIDIA requires at least N+1 power with N equal to two sources, each sized for 50% of the total peak load.
What NVIDIA says (5)
“The system includes six internal power supply units.”
“Four of the six power supplies must be energized for the system to operate.”
“Due to this requirement, the data center must minimally provide N+1 power, where N equals two power sources.”
“Each power source must be sized to support 50% of the total peak load.”
“Not acceptable for DGX H100 systems”
Concurrent maintainability means you can service any power or cooling part without shutting down the IT load. NVIDIA says the data center should generally meet or exceed Uptime Institute Tier 3, or an equivalent standard. That includes concurrent maintainability and no single point of failure.
What NVIDIA says (1)
“Generally, the data center should meet or exceed Uptime Institute Tier 3 design standards, or alternatively the TIA942-B Rated 3 or EN50600 Availability Class 3 design standards, including concurrent maintainability and no single point of failure.”
Key terms
- N+1 redundancy: Providing the N power sources you need plus one spare, so one can fail without an outage.
- Aisle containment: Enclosing the hot or cold aisle so hot exhaust air cannot mix back into server inlets.
- Gradient: The correction signal computed in training that says how much, and in which direction, to adjust each weight.
Try it
Sample question
When choosing floor space for a DGX SuperPOD, what does NVIDIA advise about the future?
Show the answer
Answer: Reserve floor space for the future state of the system, and keep SUs close together because cable lengths are limited.
A well-planned deployment thinks about growth from day one. NVIDIA says to reserve floor space for the future state of the system, not just the initial state. It also warns that performance-based cable length limits prohibit placing racks or SUs too far apart.
What NVIDIA says (2)
“A well-planned deployment will consider future expansion and reserve sufficient floor space for the future state of the system, not just the initial state.”
“Performance-based cable length limitations prohibit distributing the racks or scalable units too far from one another.”
Practice 2.6 (4 questions) Full AI Infrastructure guide
← 2.5 Components of an accelerated cluster · 2.7 Networking requirements for AI →