1.2 Managing and preprocessing multimodal data
WebDataset shards, the two NeVA data phases, precomputed encodings and GPU data loading with DALI.
Key points
Pretraining aligns image features with the language model on many caption pairs. Fine-tuning teaches the model to follow instructions about images. Each phase needs a different dataset.
What NVIDIA says (1)
“The NeVA model training encompasses two phases: pretraining and fine-tuning. Each phase mandates a unique dataset.”
WebDataset is a format that packs samples into tar archives called shards. Equal-size shards stream well during training.
What NVIDIA says (1)
“The pipeline processes the dataset into the WebDataset format, consisting of tar files of equal sizes for efficient training.”
Images can fail to download because of network errors or removed files. That leaves uneven tar files. The optional reorganize_tar stage repacks them so each holds the same number of image-text pairs.
What NVIDIA says (2)
“some images may fail to download, resulting in uneven tar files with varying number of examples each.”
“you need to re-organize the contents of the tar files so that each one contains an equal number of image-text pairs.”
A frozen encoder is one whose weights do not change in training. Its output for a given input never changes, so you can compute it once. NeMo calls this precaching the encodings. VAE means variational autoencoder.
What NVIDIA says (2)
“Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”
“Precaching these encodings can significantly enhance training throughput.”
DALI (Data Loading Library) loads and preprocesses image, video and audio data. It offloads that work from the CPU to the GPU. DALI means the NVIDIA Data Loading Library. SDK means software development kit. LLM means large language model.
What NVIDIA says (2)
“The NVIDIA Data Loading Library (DALI) is a GPU-accelerated library for data loading and pre-processing to accelerate deep learning applications.”
“DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the GPU.”
DALI is built for the media types used in multimodal training. It provides optimized building blocks for images, video and audio. DALI means the NVIDIA Data Loading Library.
What NVIDIA says (1)
“It provides a collection of highly optimized building blocks for loading and processing image, video and audio data.”
A dataset license sets how the data may be used. NVIDIA's docs say each user must review the content and license before use.
What NVIDIA says (1)
“It is the responsibility of each user to check the content of the dataset, review the applicable licenses, and determine if it is suitable for their intended use.”
Key terms
- WebDataset: A data format that packs training samples into tar files, called shards, that stream well during training.
- NVIDIA DALI: A GPU-accelerated library for loading and preprocessing image, video and audio data.
Sample question
NeVA training has two phases. What does that mean for your data?
Show the answer
Answer: Pretraining and fine-tuning each need their own dataset
Pretraining aligns image features with the language model on many caption pairs. Fine-tuning teaches the model to follow instructions about images. Each phase needs a different dataset.
What NVIDIA says (1)
“The NeVA model training encompasses two phases: pretraining and fine-tuning. Each phase mandates a unique dataset.”
Practice 1.2 (7 questions) Full Experimentation guide
← 1.1 Developing and testing multimodal models · 1.3 Explainability with multimodal models →