2.2 AI data center architecture

NCP-AIO · Administration (23% of the exam) · Official objective: “Describe data center architecture for AI Workloads”

The SuperPOD fabrics, storage, management network and Multi-Node NVLink.

Key points

  1. A fabric is a network of switches and cables. Storage reads and writes would otherwise compete with GPU-to-GPU traffic.

    What NVIDIA says (1)

    “Storage traffic is dedicated to its own fabric to remove interference with the node-to-node application traffic that can degrade overall performance.”

    — NVIDIA Mission Control Administration Guide: Overview

  2. Out-of-band (OOB) means separate from the data networks. It reaches the management controllers even when an OS is down.

    What NVIDIA says (1)

    “The out-of-band Ethernet network is used for system management using the BMC and provides connectivity to manage all networking equipment.”

    — NVIDIA Mission Control Administration Guide: Overview

  3. HSS is shared, high-bandwidth storage for every node. Home directories use a separate NFS share.

    What NVIDIA says (1)

    “High-speed storage (HSS) provides shared storage to all nodes in the DGX SuperPOD. Store datasets, checkpoints, and other large files here.”

    — NVIDIA Mission Control Administration Guide: Overview

  4. Users do not log in to GPU nodes directly. They work on login nodes, which are Slurm clients with the file systems mounted.

    What NVIDIA says (1)

    “Entry point to the DGX SuperPOD for users. CPU-based nodes that are Slurm clients with filesystems mounted to support development, job submission, job monitoring, and file management.”

    — NVIDIA Mission Control Administration Guide: Overview

  5. NVLink connects GPUs directly at high speed. On GB200 and GB300, NVLink switches extend this across trays in a rack.

    What NVIDIA says (1)

    “Multi-Node NVLink is a capability enabled over an NVLink Switch network where multiple systems are interconnected to form a large GPU memory fabric also known as an NVLink Domain.”

    — NVIDIA Mission Control Administration Guide: Overview

Key terms

Sample question

Why does the DGX SuperPOD put storage traffic on its own InfiniBand fabric?

Show the answer

Answer: To remove interference with node-to-node application traffic

A fabric is a network of switches and cables. Storage reads and writes would otherwise compete with GPU-to-GPU traffic.

What NVIDIA says (1)

“Storage traffic is dedicated to its own fabric to remove interference with the node-to-node application traffic that can degrade overall performance.”

— NVIDIA Mission Control Administration Guide: Overview

Practice 2.2 (5 questions) Full Administration guide

← 2.1 Administering Slurm · 2.3 Administering Run:ai →