6.4 Training optimization
Sequence packing, data parallelism and precomputed encodings.
Key points
Padding fills short sequences up to a fixed length, wasting compute. Sequence packing joins several examples into one long sequence instead. LLM means large language model.
What NVIDIA says (2)
“Many sequences are short, and a few are very long, conforming to Zipf’s Law.”
“Sequence packing is a training technique wherein multiple training sequences (examples) are concatenated into one long sequence (pack).”
Each GPU processes its share of the batch. Gradients are then synchronized so the copies stay the same.
What NVIDIA says (1)
“Data Parallelism (DP) replicates the model across multiple GPUs. Data batches are evenly distributed between GPUs and the data-parallel GPUs process them independently.”
Frozen encoders always give the same output, so you can compute it once and reuse it. VAE means variational autoencoder.
What NVIDIA says (1)
“Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”
Key terms
- Sequence packing: Joining several short training examples into one long sequence to avoid padding.
- Data parallelism: Copying a model to several GPUs and splitting each batch between them.
Sample question
Multimodal LLM datasets have many short and a few very long sequences. Which NeMo technique avoids wasted padding?
Show the answer
Answer: Sequence packing
Padding fills short sequences up to a fixed length, wasting compute. Sequence packing joins several examples into one long sequence instead. LLM means large language model.
What NVIDIA says (2)
“Many sequences are short, and a few are very long, conforming to Zipf’s Law.”
“Sequence packing is a training technique wherein multiple training sequences (examples) are concatenated into one long sequence (pack).”
Practice 6.4 (3 questions) Full Performance Optimization guide
← 6.3 Transfer learning for performance · 7.1 Ethical principles of trustworthy AI →