4.13 NeMo burn-in

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Perform NeMo burn-in.”

Using NeMo training runs and NVIDIA recipes as an end-to-end load test.

Key points

  1. NeMo Framework is NVIDIA's framework for training and deploying large language models. A real training job is a good burn-in, because it loads GPUs, memory, network and storage together. NVIDIA's example uses synthetic data, so the first run does not depend on preparing a dataset.

    What NVIDIA says (2)

    “The document walks through using a DGX Cloud Slurm cluster as a user to launch a simple pretraining job, targeting synthetic data to minimize dependencies for initial use.”

    — DGX Cloud Slurm: Workload Examples

    “In order to pull the NeMo FW training container from NGC, the previously noted NGC API key needs to be added to a configuration file in the DGX Cloud Slurm cluster.”

    — DGX Cloud Slurm: Workload Examples

  2. A performance recipe is a packaged, repeatable benchmark run, such as NeMo pretraining. Each one is measured on NVIDIA reference architectures to set a baseline. Run the small system info recipe first. It catches setup problems before you spend hours of GPU time.

    What NVIDIA says (2)

    “we recommend running the system info recipe to collect basic system information and check for common cluster configuration issues”

    — NVIDIA Exemplar Performance: Performance Recipes README

    “These workloads are tested against NVIDIA Reference Architectures to establish baselines for comparison.”

    — NVIDIA Exemplar Performance: Performance Recipes README

Key terms

Sample question

For a first NeMo training run on a new Slurm cluster, what data does NVIDIA's DGX Cloud example use, and why?

Show the answer

Answer: Synthetic data, to minimize dependencies.

NeMo Framework is NVIDIA's framework for training and deploying large language models. A real training job is a good burn-in, because it loads GPUs, memory, network and storage together. NVIDIA's example uses synthetic data, so the first run does not depend on preparing a dataset.

What NVIDIA says (2)

“The document walks through using a DGX Cloud Slurm cluster as a user to launch a simple pretraining job, targeting synthetic data to minimize dependencies for initial use.”

— DGX Cloud Slurm: Workload Examples

“In order to pull the NeMo FW training container from NGC, the previously noted NGC API key needs to be added to a configuration file in the DGX Cloud Slurm cluster.”

— DGX Cloud Slurm: Workload Examples

Practice 4.13 (2 questions) Full Cluster Test and Verification guide

← 4.12 HPL burn-in · 4.14 Testing storage →