2.6 Facility requirements

NCA-AIIO · AI Infrastructure (40% of the exam) · Official objective: “Identify facility requirements”

Redundancy standards, power feeds, floor space and airflow for an AI data center.

Key points

  1. A well-planned deployment thinks about growth from day one. NVIDIA says to reserve floor space for the future state of the system, not just the initial state. It also warns that performance-based cable length limits prohibit placing racks or SUs too far apart.

    What NVIDIA says (2)

    “A well-planned deployment will consider future expansion and reserve sufficient floor space for the future state of the system, not just the initial state.”

    — DGX SuperPOD Data Center Design (H100): Infrastructure

    “Performance-based cable length limitations prohibit distributing the racks or scalable units too far from one another.”

    — DGX SuperPOD Data Center Design (H100): Infrastructure

  2. A resource is oversubscribed when demand exceeds capacity. NVIDIA says oversubscribed cooling can be mitigated by lowering rack density or by spacing racks apart. In that case, the cooling shortfall is 'paid for' with floor space. It also stresses first optimizing airflow, for example with aisle containment, which stops hot exhaust air from recirculating to server inlets.

    What NVIDIA says (3)

    “Oversubscribed cooling can sometimes be mitigated by lowering the rack density thereby reducing the cooling demand per rack footprint, or by spacing the racks further apart to aggregate the cooling capacity of more than one rack footprint to each populated rack.”

    — DGX SuperPOD Data Center Design (H100): Cooling

    “In this example, the resource constraint in cooling is “paid for” using floor space.”

    — DGX SuperPOD Data Center Design (H100): Planning

    “In either case, the main benefit of aisle containment is the prevention of air recirculation from the hot aisle to the cold aisle, which artificially increases the supply air temperature at the inlet of the servers, significantly reducing its heat exchange potential.”

    — DGX SuperPOD Data Center Design (H100): Cooling

  3. A PSU (power supply unit) converts facility power for the server. NVIDIA says four of the six PSUs must be energized for the system to operate. With only two feeds, losing one would drop a system to three PSUs and shut it down. So NVIDIA requires at least N+1 power with N equal to two sources, each sized for 50% of the total peak load.

    What NVIDIA says (5)

    “The system includes six internal power supply units.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Four of the six power supplies must be energized for the system to operate.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Due to this requirement, the data center must minimally provide N+1 power, where N equals two power sources.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Each power source must be sized to support 50% of the total peak load.”

    — DGX SuperPOD Data Center Design (H100): Electrical

    “Not acceptable for DGX H100 systems”

    — DGX SuperPOD Data Center Design (H100): Electrical

  4. Concurrent maintainability means you can service any power or cooling part without shutting down the IT load. NVIDIA says the data center should generally meet or exceed Uptime Institute Tier 3, or an equivalent standard. That includes concurrent maintainability and no single point of failure.

    What NVIDIA says (1)

    “Generally, the data center should meet or exceed Uptime Institute Tier 3 design standards, or alternatively the TIA942-B Rated 3 or EN50600 Availability Class 3 design standards, including concurrent maintainability and no single point of failure.”

    — DGX SuperPOD Data Center Design (H100): Electrical

Key terms

Try it

Sample question

When choosing floor space for a DGX SuperPOD, what does NVIDIA advise about the future?

Show the answer

Answer: Reserve floor space for the future state of the system, and keep SUs close together because cable lengths are limited.

A well-planned deployment thinks about growth from day one. NVIDIA says to reserve floor space for the future state of the system, not just the initial state. It also warns that performance-based cable length limits prohibit placing racks or SUs too far apart.

What NVIDIA says (2)

“A well-planned deployment will consider future expansion and reserve sufficient floor space for the future state of the system, not just the initial state.”

— DGX SuperPOD Data Center Design (H100): Infrastructure

“Performance-based cable length limitations prohibit distributing the racks or scalable units too far from one another.”

— DGX SuperPOD Data Center Design (H100): Infrastructure

Practice 2.6 (4 questions) Full AI Infrastructure guide

← 2.5 Components of an accelerated cluster · 2.7 Networking requirements for AI →