1.1 Developing and testing multimodal models
How NeVA, CLIP, Stable Diffusion, Video NeVA and SpeechLLMs are put together, and what to check when you build one.
Key points
A multimodal model works with more than one kind of data, such as images and text. NeVA (NeMo Vision and Language Assistant) joins a large language model (LLM) to a vision encoder. A vision encoder is a network that turns an image into feature vectors.
What NVIDIA says (1)
“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”
The vision encoder in NeVA is the pretrained CLIP ViT-L/14. A projection matrix is a learned layer that changes vectors from one space into another. NeVA uses it to blend visual features with the language embeddings. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (2)
“NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”
“The encoder retrieves visual features from images and intertwines them with language embeddings using a modifiable projection matrix.”
CLIP (Contrastive Language-Image Pre-training) learns from image-caption pairs. Contrastive training pulls matching pairs together and pushes mismatched pairs apart. The result is a shared space where images and text can be compared.
What NVIDIA says (2)
“The essence of CLIP is to train both an image encoder and a text encoder from scratch.”
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
Stable Diffusion generates images from text. The U-Net predicts noise. The variational autoencoder (VAE) compresses images into a smaller latent space. The CLIP text encoder turns the prompt into embeddings that guide the U-Net. CLIP means Contrastive Language-Image Pre-training. GAN means generative adversarial network.
What NVIDIA says (1)
“Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”
A modality is a type of data, such as text, image, audio or video. Video NeVA treats a video as a series of frames. Its config sets how many frames to take, for example num_frames.
What NVIDIA says (1)
“Video NeVa adds support for video modality in NeVa by representing video as multiple image frames.”
A SpeechLLM is a large language model that also accepts audio. An audio encoder turns sound into audio embeddings. The modality adapter maps those into the LLM's embedding space so the LLM can use them with the text prompt.
What NVIDIA says (1)
“A modality adapter that processes the audio embeddings and produces a sequence of embeddings in the same latent space as the token embeddings of a pretrained LLM.”
Concatenation places the speech embeddings next to the text embeddings in time. Cross-attention lets one sequence (text) look up information in another (speech). NeMo adds the cross-attention module only before the LLM to keep cost down. LLM means large language model.
What NVIDIA says (3)
“One way to incorporate speech into an LLM is to concatenate speech features with the token embeddings of the input text prompt before feeding them into the LLM.”
“The Speech-Augmented Language Model (SALM) follows this approach.”
“Another approach is to use a cross-attention mechanism, where text embeddings attend to speech embeddings to extract task-specific information.”
The projection maps image features into the language model's embedding space. A multilayer perceptron (MLP) is a small stack of fully connected layers. A two-layer MLP is more expressive than one linear layer.
What NVIDIA says (1)
“Transitioning from a linear to a dual-layer MLP projection markedly bolsters LLaVA-1.5’s multimodal faculties”
Key terms
- Multimodal model: A model that works with more than one type of data, such as text, images, audio or video.
- Modality: One type of data, such as text, image, audio or video.
- NeVA: NVIDIA's NeMo vision-language model that joins a large language model to a CLIP vision encoder through a projection layer.
- Vision encoder: A network that turns an image into feature vectors that other parts of a model can use.
- Projection layer: A learned layer that maps vectors from one embedding space into another, such as image features into text embeddings.
- CLIP: A model with an image encoder and a text encoder trained together so matching images and captions land close in one vector space.
- Stable Diffusion: A latent text-to-image diffusion model made of a U-Net, a VAE image encoder and a CLIP text encoder.
- SpeechLLM: A large language model that also takes audio, through an audio encoder and a modality adapter.
- Modality adapter: A module that maps embeddings from one modality into the embedding space of a language model.
- Cross-attention: Attention in which one sequence, such as text or image features, looks up information in another sequence.
Try it
Sample question
NeVA is NVIDIA's vision-language model in NeMo. What does it combine?
Show the answer
Answer: A large language model with a vision encoder
A multimodal model works with more than one kind of data, such as images and text. NeVA (NeMo Vision and Language Assistant) joins a large language model (LLM) to a vision encoder. A vision encoder is a network that turns an image into feature vectors.
What NVIDIA says (1)
“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”