6.4 Training optimization

NCA-GENM · Performance Optimization (10% of the exam) · Official objective: “Assist in model training and training optimization under the supervision of a senior team member.”

Sequence packing, data parallelism and precomputed encodings.

Key points

  1. Padding fills short sequences up to a fixed length, wasting compute. Sequence packing joins several examples into one long sequence instead. LLM means large language model.

    What NVIDIA says (2)

    “Many sequences are short, and a few are very long, conforming to Zipf’s Law.”

    — NeMo Framework 24.09: Sequence Packing for NeVA

    “Sequence packing is a training technique wherein multiple training sequences (examples) are concatenated into one long sequence (pack).”

    — NeMo Framework 24.09: Sequence Packing for NeVA

  2. Each GPU processes its share of the batch. Gradients are then synchronized so the copies stay the same.

    What NVIDIA says (1)

    “Data Parallelism (DP) replicates the model across multiple GPUs. Data batches are evenly distributed between GPUs and the data-parallel GPUs process them independently.”

    — NeMo Framework 24.09: Parallelisms

  3. Frozen encoders always give the same output, so you can compute it once and reuse it. VAE means variational autoencoder.

    What NVIDIA says (1)

    “Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”

    — NeMo Framework 24.09: Stable Diffusion

Key terms

Sample question

Multimodal LLM datasets have many short and a few very long sequences. Which NeMo technique avoids wasted padding?

Show the answer

Answer: Sequence packing

Padding fills short sequences up to a fixed length, wasting compute. Sequence packing joins several examples into one long sequence instead. LLM means large language model.

What NVIDIA says (2)

“Many sequences are short, and a few are very long, conforming to Zipf’s Law.”

— NeMo Framework 24.09: Sequence Packing for NeVA

“Sequence packing is a training technique wherein multiple training sequences (examples) are concatenated into one long sequence (pack).”

— NeMo Framework 24.09: Sequence Packing for NeVA

Practice 6.4 (3 questions) Full Performance Optimization guide

← 6.3 Transfer learning for performance · 7.1 Ethical principles of trustworthy AI →