1.2 Managing and preprocessing multimodal data

NCA-GENM · Experimentation (25% of the exam) · Official objective: “Manage and preprocess data from various sources.”

WebDataset shards, the two NeVA data phases, precomputed encodings and GPU data loading with DALI.

Key points

  1. Pretraining aligns image features with the language model on many caption pairs. Fine-tuning teaches the model to follow instructions about images. Each phase needs a different dataset.

    What NVIDIA says (1)

    “The NeVA model training encompasses two phases: pretraining and fine-tuning. Each phase mandates a unique dataset.”

    — NeMo Framework 24.09: Multimodal Language Model Datasets

  2. WebDataset is a format that packs samples into tar archives called shards. Equal-size shards stream well during training.

    What NVIDIA says (1)

    “The pipeline processes the dataset into the WebDataset format, consisting of tar files of equal sizes for efficient training.”

    — NeMo Framework 24.09: Text to Image Datasets

  3. Images can fail to download because of network errors or removed files. That leaves uneven tar files. The optional reorganize_tar stage repacks them so each holds the same number of image-text pairs.

    What NVIDIA says (2)

    “some images may fail to download, resulting in uneven tar files with varying number of examples each.”

    — NeMo Framework 24.09: Text to Image Datasets

    “you need to re-organize the contents of the tar files so that each one contains an equal number of image-text pairs.”

    — NeMo Framework 24.09: Text to Image Datasets

  4. A frozen encoder is one whose weights do not change in training. Its output for a given input never changes, so you can compute it once. NeMo calls this precaching the encodings. VAE means variational autoencoder.

    What NVIDIA says (2)

    “Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”

    — NeMo Framework 24.09: Stable Diffusion

    “Precaching these encodings can significantly enhance training throughput.”

    — NeMo Framework 24.09: Text to Image Datasets

  5. DALI (Data Loading Library) loads and preprocesses image, video and audio data. It offloads that work from the CPU to the GPU. DALI means the NVIDIA Data Loading Library. SDK means software development kit. LLM means large language model.

    What NVIDIA says (2)

    “The NVIDIA Data Loading Library (DALI) is a GPU-accelerated library for data loading and pre-processing to accelerate deep learning applications.”

    — NVIDIA DALI User Guide

    “DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the GPU.”

    — NVIDIA DALI User Guide

  6. DALI is built for the media types used in multimodal training. It provides optimized building blocks for images, video and audio. DALI means the NVIDIA Data Loading Library.

    What NVIDIA says (1)

    “It provides a collection of highly optimized building blocks for loading and processing image, video and audio data.”

    — NVIDIA DALI User Guide

  7. A dataset license sets how the data may be used. NVIDIA's docs say each user must review the content and license before use.

    What NVIDIA says (1)

    “It is the responsibility of each user to check the content of the dataset, review the applicable licenses, and determine if it is suitable for their intended use.”

    — NeMo Framework 24.09: Text to Image Datasets

Key terms

Sample question

NeVA training has two phases. What does that mean for your data?

Show the answer

Answer: Pretraining and fine-tuning each need their own dataset

Pretraining aligns image features with the language model on many caption pairs. Fine-tuning teaches the model to follow instructions about images. Each phase needs a different dataset.

What NVIDIA says (1)

“The NeVA model training encompasses two phases: pretraining and fine-tuning. Each phase mandates a unique dataset.”

— NeMo Framework 24.09: Multimodal Language Model Datasets

Practice 1.2 (7 questions) Full Experimentation guide

← 1.1 Developing and testing multimodal models · 1.3 Explainability with multimodal models →