1.1 Developing and testing multimodal models

NCA-GENM · Experimentation (25% of the exam) · Official objective: “Assist in developing and testing multimodal AI models.”

How NeVA, CLIP, Stable Diffusion, Video NeVA and SpeechLLMs are put together, and what to check when you build one.

Key points

  1. A multimodal model works with more than one kind of data, such as images and text. NeVA (NeMo Vision and Language Assistant) joins a large language model (LLM) to a vision encoder. A vision encoder is a network that turns an image into feature vectors.

    What NVIDIA says (1)

    “It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”

    — NeMo Framework 24.09: NeVA

  2. The vision encoder in NeVA is the pretrained CLIP ViT-L/14. A projection matrix is a learned layer that changes vectors from one space into another. NeVA uses it to blend visual features with the language embeddings. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (2)

    “NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”

    — NeMo Framework 24.09: NeVA

    “The encoder retrieves visual features from images and intertwines them with language embeddings using a modifiable projection matrix.”

    — NeMo Framework 24.09: NeVA

  3. CLIP (Contrastive Language-Image Pre-training) learns from image-caption pairs. Contrastive training pulls matching pairs together and pushes mismatched pairs apart. The result is a shared space where images and text can be compared.

    What NVIDIA says (2)

    “The essence of CLIP is to train both an image encoder and a text encoder from scratch.”

    — NeMo Framework 24.09: CLIP

    “maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

    — NeMo Framework 24.09: CLIP

  4. Stable Diffusion generates images from text. The U-Net predicts noise. The variational autoencoder (VAE) compresses images into a smaller latent space. The CLIP text encoder turns the prompt into embeddings that guide the U-Net. CLIP means Contrastive Language-Image Pre-training. GAN means generative adversarial network.

    What NVIDIA says (1)

    “Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”

    — NeMo Framework 24.09: Stable Diffusion

  5. A modality is a type of data, such as text, image, audio or video. Video NeVA treats a video as a series of frames. Its config sets how many frames to take, for example num_frames.

    What NVIDIA says (1)

    “Video NeVa adds support for video modality in NeVa by representing video as multiple image frames.”

    — NeMo Framework 24.09: Video NeVA

  6. A SpeechLLM is a large language model that also accepts audio. An audio encoder turns sound into audio embeddings. The modality adapter maps those into the LLM's embedding space so the LLM can use them with the text prompt.

    What NVIDIA says (1)

    “A modality adapter that processes the audio embeddings and produces a sequence of embeddings in the same latent space as the token embeddings of a pretrained LLM.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

  7. Concatenation places the speech embeddings next to the text embeddings in time. Cross-attention lets one sequence (text) look up information in another (speech). NeMo adds the cross-attention module only before the LLM to keep cost down. LLM means large language model.

    What NVIDIA says (3)

    “One way to incorporate speech into an LLM is to concatenate speech features with the token embeddings of the input text prompt before feeding them into the LLM.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

    “The Speech-Augmented Language Model (SALM) follows this approach.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

    “Another approach is to use a cross-attention mechanism, where text embeddings attend to speech embeddings to extract task-specific information.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

  8. The projection maps image features into the language model's embedding space. A multilayer perceptron (MLP) is a small stack of fully connected layers. A two-layer MLP is more expressive than one linear layer.

    What NVIDIA says (1)

    “Transitioning from a linear to a dual-layer MLP projection markedly bolsters LLaVA-1.5’s multimodal faculties”

    — NeMo Framework 24.09: NeVA

Key terms

Try it

Sample question

NeVA is NVIDIA's vision-language model in NeMo. What does it combine?

Show the answer

Answer: A large language model with a vision encoder

A multimodal model works with more than one kind of data, such as images and text. NeVA (NeMo Vision and Language Assistant) joins a large language model (LLM) to a vision encoder. A vision encoder is a network that turns an image into feature vectors.

What NVIDIA says (1)

“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”

— NeMo Framework 24.09: NeVA

Practice 1.1 (8 questions) Full Experimentation guide

1.2 Managing and preprocessing multimodal data →