2.5 Components of an accelerated cluster
Compute, networking, management and storage, and how training uses storage.
Key points
I/O (input/output) here means reading and writing data to storage. Training passes over the same dataset many times. Each pass is an epoch. NVIDIA says the key I/O operation in DL training is re-read. That is why caching the data after the first read helps so much.
What NVIDIA says (2)
“The key I/O operation in DL training is re-read.”
“It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”
DMA (direct memory access) lets a device move data without the CPU copying each byte. A bounce buffer is an extra copy in CPU memory on the way to the GPU. GDS enables a direct DMA path between GPU memory and storage that avoids that bounce buffer. NVIDIA says this can relieve bandwidth bottlenecks and lower latency and CPU load.
What NVIDIA says (2)
“GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which avoids a bounce buffer through the CPU.”
“Using this direct path can relieve system bandwidth bottlenecks and decrease the latency and utilization load on the CPU.”
A checkpoint is a saved copy of the model's state during training. If a node fails, the job restarts from the last checkpoint instead of from the beginning. NVIDIA says writing checkpoints is necessary for fault tolerance as models grow, and that write performance can also be important.
What NVIDIA says (2)
“As DL models grow in size and time-to-train, writing checkpoints is necessary for fault tolerance.”
“Write performance can also be important.”
An accelerated cluster is more than its GPU servers. NVIDIA defines the DGX SuperPOD architecture as a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage. Each part has to keep up with the others or the GPUs sit idle.
What NVIDIA says (1)
“The DGX SuperPOD architecture is a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage.”
Key terms
- DGX SuperPOD: NVIDIA's reference design for an AI cluster made of DGX systems, InfiniBand and Ethernet networking, management nodes and storage.
- GPUDirect Storage: A direct path that moves data between storage and GPU memory without a copy through CPU memory.
- Checkpoint: A saved copy of training state so a job can resume after a failure.
- NVLink: NVIDIA's direct GPU-to-GPU interconnect inside a server.
- PCI Express: The standard interconnect inside a server that links the CPU to GPUs, network adapters and storage.
- Bandwidth: How much data a link or memory can move per second, for example in GB/s.
Try it
Sample question
What is the key storage I/O pattern in deep learning (DL) training, according to the DGX SuperPOD reference architecture?
Show the answer
Answer: Re-read: the same data is read again and again because training is iterative.
I/O (input/output) here means reading and writing data to storage. Training passes over the same dataset many times. Each pass is an epoch. NVIDIA says the key I/O operation in DL training is re-read. That is why caching the data after the first read helps so much.
What NVIDIA says (2)
“The key I/O operation in DL training is re-read.”
“It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”