4.13 NeMo burn-in
Using NeMo training runs and NVIDIA recipes as an end-to-end load test.
Key points
NeMo Framework is NVIDIA's framework for training and deploying large language models. A real training job is a good burn-in, because it loads GPUs, memory, network and storage together. NVIDIA's example uses synthetic data, so the first run does not depend on preparing a dataset.
What NVIDIA says (2)
“The document walks through using a DGX Cloud Slurm cluster as a user to launch a simple pretraining job, targeting synthetic data to minimize dependencies for initial use.”
“In order to pull the NeMo FW training container from NGC, the previously noted NGC API key needs to be added to a configuration file in the DGX Cloud Slurm cluster.”
A performance recipe is a packaged, repeatable benchmark run, such as NeMo pretraining. Each one is measured on NVIDIA reference architectures to set a baseline. Run the small system info recipe first. It catches setup problems before you spend hours of GPU time.
What NVIDIA says (2)
“we recommend running the system info recipe to collect basic system information and check for common cluster configuration issues”
“These workloads are tested against NVIDIA Reference Architectures to establish baselines for comparison.”
Key terms
- Burn-in: Running a heavy workload for a long time to expose weak parts before production.
Sample question
For a first NeMo training run on a new Slurm cluster, what data does NVIDIA's DGX Cloud example use, and why?
Show the answer
Answer: Synthetic data, to minimize dependencies.
NeMo Framework is NVIDIA's framework for training and deploying large language models. A real training job is a good burn-in, because it loads GPUs, memory, network and storage together. NVIDIA's example uses synthetic data, so the first run does not depend on preparing a dataset.
What NVIDIA says (2)
“The document walks through using a DGX Cloud Slurm cluster as a user to launch a simple pretraining job, targeting synthetic data to minimize dependencies for initial use.”
“In order to pull the NeMo FW training container from NGC, the previously noted NGC API key needs to be added to a configuration file in the DGX Cloud Slurm cluster.”
Practice 4.13 (2 questions) Full Cluster Test and Verification guide