2.1 Hardware for training workloads
What a training node provides: GPUs, GPU memory, local NVMe and fast networking.
Key points
A cache is fast, nearby storage that keeps a copy of data so it does not have to be fetched again over the network. The DGX SuperPOD reference architecture rates storage needs by workload. It puts NLP at the 'Good' level because NLP datasets generally fit within the local cache. Larger datasets, such as compressed images, push the requirement higher.
What NVIDIA says (2)
“Good Natural Language Processing (NLP) Datasets generally fit within local cache”
“Ideally, data is cached during the first read of the dataset, so data does not have to be retrieved across the network.”
GPU memory is the memory on the GPU itself. Model weights and intermediate values must fit in it during training. The DGX H100 user guide lists 8 x NVIDIA H100 GPUs that provide 640 GB total GPU memory. That is 80 GB per GPU. If a model needs more than that, it is split across GPUs, which is why the GPU-to-GPU link matters.
What NVIDIA says (2)
“The DGX H100/H200 systems are built on eight NVIDIA H100 Tensor Core GPUs or eight NVIDIA H200 Tensor Core GPUs.”
“For H100: 8 x NVIDIA H100 GPUs that provide 640 GB total GPU memory”
NVMe (Non-Volatile Memory Express) is a fast protocol for solid-state drives. A DGX H100 has a RAID 0 array of NVMe U.2 drives labeled as the data cache. The reference architecture says this local NVMe storage can be used for caching or staging data. Staging means copying data close to the GPUs before training starts.
What NVIDIA says (2)
“In addition, the DGX H100 system provides local NVMe storage that can also be used for caching or staging data.”
“Storage (Data Cache) 8 x 3.84 TB NVMe U.2 SED (ea) in RAID 0 array”
Try it
Sample question
You are sizing storage for a natural language processing (NLP) training job on DGX H100 systems. What does the DGX SuperPOD reference architecture say about NLP datasets and the local cache?
Show the answer
Answer: NLP datasets generally fit within the local cache, so they sit at the 'Good' storage performance level.
A cache is fast, nearby storage that keeps a copy of data so it does not have to be fetched again over the network. The DGX SuperPOD reference architecture rates storage needs by workload. It puts NLP at the 'Good' level because NLP datasets generally fit within the local cache. Larger datasets, such as compressed images, push the requirement higher.
What NVIDIA says (2)
“Good Natural Language Processing (NLP) Datasets generally fit within local cache”
“Ideally, data is cached during the first read of the dataset, so data does not have to be retrieved across the network.”
Practice 2.1 (3 questions) Full AI Infrastructure guide
← 1.8 GPU vs CPU architecture · 2.2 Scaling GPU infrastructure →