2.5 Components of an accelerated cluster

NCA-AIIO · AI Infrastructure (40% of the exam) · Official objective: “Identify key components and considerations of a cluster of an accelerated infrastructure”

Compute, networking, management and storage, and how training uses storage.

Key points

  1. I/O (input/output) here means reading and writing data to storage. Training passes over the same dataset many times. Each pass is an epoch. NVIDIA says the key I/O operation in DL training is re-read. That is why caching the data after the first read helps so much.

    What NVIDIA says (2)

    “The key I/O operation in DL training is re-read.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

    “It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

  2. DMA (direct memory access) lets a device move data without the CPU copying each byte. A bounce buffer is an extra copy in CPU memory on the way to the GPU. GDS enables a direct DMA path between GPU memory and storage that avoids that bounce buffer. NVIDIA says this can relieve bandwidth bottlenecks and lower latency and CPU load.

    What NVIDIA says (2)

    “GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which avoids a bounce buffer through the CPU.”

    — GPUDirect Storage Overview

    “Using this direct path can relieve system bandwidth bottlenecks and decrease the latency and utilization load on the CPU.”

    — GPUDirect Storage Overview

  3. A checkpoint is a saved copy of the model's state during training. If a node fails, the job restarts from the last checkpoint instead of from the beginning. NVIDIA says writing checkpoints is necessary for fault tolerance as models grow, and that write performance can also be important.

    What NVIDIA says (2)

    “As DL models grow in size and time-to-train, writing checkpoints is necessary for fault tolerance.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

    “Write performance can also be important.”

    — DGX SuperPOD H100 Reference Architecture: Storage Architecture

  4. An accelerated cluster is more than its GPU servers. NVIDIA defines the DGX SuperPOD architecture as a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage. Each part has to keep up with the others or the GPUs sit idle.

    What NVIDIA says (1)

    “The DGX SuperPOD architecture is a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage.”

    — DGX SuperPOD H100 Reference Architecture: Architecture

Key terms

Try it

Sample question

What is the key storage I/O pattern in deep learning (DL) training, according to the DGX SuperPOD reference architecture?

Show the answer

Answer: Re-read: the same data is read again and again because training is iterative.

I/O (input/output) here means reading and writing data to storage. Training passes over the same dataset many times. Each pass is an epoch. NVIDIA says the key I/O operation in DL training is re-read. That is why caching the data after the first read helps so much.

What NVIDIA says (2)

“The key I/O operation in DL training is re-read.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

“It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

Practice 2.5 (4 questions) Full AI Infrastructure guide

← 2.4 On-prem vs cloud · 2.6 Facility requirements →