Experimentation
25% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI
1.1 Developing and testing multimodal models
How NeVA, CLIP, Stable Diffusion, Video NeVA and SpeechLLMs are put together, and what to check when you build one.
Key points
A multimodal model works with more than one kind of data, such as images and text. NeVA (NeMo Vision and Language Assistant) joins a large language model (LLM) to a vision encoder. A vision encoder is a network that turns an image into feature vectors.
What NVIDIA says (1)
“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”
The vision encoder in NeVA is the pretrained CLIP ViT-L/14. A projection matrix is a learned layer that changes vectors from one space into another. NeVA uses it to blend visual features with the language embeddings. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (2)
“NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”
“The encoder retrieves visual features from images and intertwines them with language embeddings using a modifiable projection matrix.”
CLIP (Contrastive Language-Image Pre-training) learns from image-caption pairs. Contrastive training pulls matching pairs together and pushes mismatched pairs apart. The result is a shared space where images and text can be compared.
What NVIDIA says (2)
“The essence of CLIP is to train both an image encoder and a text encoder from scratch.”
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
Stable Diffusion generates images from text. The U-Net predicts noise. The variational autoencoder (VAE) compresses images into a smaller latent space. The CLIP text encoder turns the prompt into embeddings that guide the U-Net. CLIP means Contrastive Language-Image Pre-training. GAN means generative adversarial network.
What NVIDIA says (1)
“Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”
A modality is a type of data, such as text, image, audio or video. Video NeVA treats a video as a series of frames. Its config sets how many frames to take, for example num_frames.
What NVIDIA says (1)
“Video NeVa adds support for video modality in NeVa by representing video as multiple image frames.”
A SpeechLLM is a large language model that also accepts audio. An audio encoder turns sound into audio embeddings. The modality adapter maps those into the LLM's embedding space so the LLM can use them with the text prompt.
What NVIDIA says (1)
“A modality adapter that processes the audio embeddings and produces a sequence of embeddings in the same latent space as the token embeddings of a pretrained LLM.”
Concatenation places the speech embeddings next to the text embeddings in time. Cross-attention lets one sequence (text) look up information in another (speech). NeMo adds the cross-attention module only before the LLM to keep cost down. LLM means large language model.
What NVIDIA says (3)
“One way to incorporate speech into an LLM is to concatenate speech features with the token embeddings of the input text prompt before feeding them into the LLM.”
“The Speech-Augmented Language Model (SALM) follows this approach.”
“Another approach is to use a cross-attention mechanism, where text embeddings attend to speech embeddings to extract task-specific information.”
The projection maps image features into the language model's embedding space. A multilayer perceptron (MLP) is a small stack of fully connected layers. A two-layer MLP is more expressive than one linear layer.
What NVIDIA says (1)
“Transitioning from a linear to a dual-layer MLP projection markedly bolsters LLaVA-1.5’s multimodal faculties”
Key terms: Multimodal model Modality NeVA Vision encoder Projection layer CLIP Stable Diffusion SpeechLLM Modality adapter Cross-attention
Try it: CLIP lab
1.2 Managing and preprocessing multimodal data
WebDataset shards, the two NeVA data phases, precomputed encodings and GPU data loading with DALI.
Key points
Pretraining aligns image features with the language model on many caption pairs. Fine-tuning teaches the model to follow instructions about images. Each phase needs a different dataset.
What NVIDIA says (1)
“The NeVA model training encompasses two phases: pretraining and fine-tuning. Each phase mandates a unique dataset.”
WebDataset is a format that packs samples into tar archives called shards. Equal-size shards stream well during training.
What NVIDIA says (1)
“The pipeline processes the dataset into the WebDataset format, consisting of tar files of equal sizes for efficient training.”
Images can fail to download because of network errors or removed files. That leaves uneven tar files. The optional reorganize_tar stage repacks them so each holds the same number of image-text pairs.
What NVIDIA says (2)
“some images may fail to download, resulting in uneven tar files with varying number of examples each.”
“you need to re-organize the contents of the tar files so that each one contains an equal number of image-text pairs.”
A frozen encoder is one whose weights do not change in training. Its output for a given input never changes, so you can compute it once. NeMo calls this precaching the encodings. VAE means variational autoencoder.
What NVIDIA says (2)
“Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”
“Precaching these encodings can significantly enhance training throughput.”
DALI (Data Loading Library) loads and preprocesses image, video and audio data. It offloads that work from the CPU to the GPU. DALI means the NVIDIA Data Loading Library. SDK means software development kit. LLM means large language model.
What NVIDIA says (2)
“The NVIDIA Data Loading Library (DALI) is a GPU-accelerated library for data loading and pre-processing to accelerate deep learning applications.”
“DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the GPU.”
DALI is built for the media types used in multimodal training. It provides optimized building blocks for images, video and audio. DALI means the NVIDIA Data Loading Library.
What NVIDIA says (1)
“It provides a collection of highly optimized building blocks for loading and processing image, video and audio data.”
A dataset license sets how the data may be used. NVIDIA's docs say each user must review the content and license before use.
What NVIDIA says (1)
“It is the responsibility of each user to check the content of the dataset, review the applicable licenses, and determine if it is suitable for their intended use.”
Key terms: WebDataset NVIDIA DALI
1.3 Explainability with multimodal models
Attention, LIME and SHAP, grounding in RAG, model cards and proxy models.
Key points
Explainable AI (XAI) is a set of tools and techniques that help people understand why a model decided something. For images, audio and text, attention can be visualized to show which parts of the input mattered.
What NVIDIA says (2)
“is a set of tools and techniques used by organizations to help people better understand why a model makes certain decisions and how it works.”
“For some data — images, audio and text — similar results can be visualized through the use of”
LIME and SHAP assign credit to each input feature for one prediction. NVIDIA notes they give literal mathematical answers that can be shown to many audiences.
What NVIDIA says (1)
“Techniques with names like LIME and SHAP offer very literal mathematical answers to this question”
Retrieval-augmented generation (RAG) fetches relevant data and gives it to the model as context. Grounding means converting other modalities into one primary modality, here text. The text descriptions also make the retrieved evidence readable by people.
What NVIDIA says (2)
“The key benefit here is that the metadata generated from the information-rich image is extremely helpful in answering objective questions.”
“The key disadvantages are preprocessing costs and losing some nuance from the image.”
RAG retrieves documents and passes them to the model. The model can then cite them, so users can verify each claim. A multimodal RAG app can cite retrieved images and charts the same way.
What NVIDIA says (1)
“Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”
A model card is a short document that describes a model, its uses and its limits. Model Card++ adds four subsections on trust topics.
What NVIDIA says (1)
“Four subsections detailing model-specific information concerning Bias, Explainability, Privacy, and Safety and Security.”
Proxy modeling uses a simple model to approximate a complex one. It gives a sense of the whole model, but it is only an approximation.
What NVIDIA says (1)
“simpler, more easily comprehended models like decision trees can be used to approximately describe the more detailed AI model.”
Pedigree is the history of how a model was made. Knowing it helps people judge when the model's outputs make sense, including for multimodal models trained on images and text.
What NVIDIA says (1)
“Explaining the pedigree of the model: How was the model trained? What data was used? How was the impact of any bias in the training data measured and mitigated?”
Key terms: Attention map Explainable AI Model card Retrieval-augmented generation
1.4 Testing multimodal data quality
NeMo Curator aesthetic and NSFW scores, semantic deduplication, heuristic filters and PII removal.
Key points
NeMo Curator is NVIDIA's data curation library. Its aesthetic classifier estimates the subjective quality of an image. A higher score means a more pleasing image. PII means personal identifiable information. NSFW means not safe for work.
What NVIDIA says (2)
“Aesthetic classifiers can be used to assess the subjective quality of an image.”
“outputs a score from 0-10 where 10 is aesthetically pleasing.”
An embedding is a vector that represents an input. The aesthetic classifier is a small linear model on top of CLIP image embeddings. CLIP means Contrastive Language-Image Pre-training. PII means personal identifiable information.
What NVIDIA says (1)
“a linear classifier that takes OpenAI CLIP ViT-L/14 image embeddings as input.”
NSFW means not safe for work. Removing unsafe content is a common step in generative AI data pipelines. PII means personal identifiable information.
What NVIDIA says (2)
“a value between 0 and 1 where 1 means the content is NSFW.”
“Removing unsafe content is common in most data processing pipelines to prevent your generative AI model from learning to produce unsafe material.”
Semantic duplicates are pairs that mean almost the same thing without being identical. Semantic deduplication embeds each sample, clusters the embeddings with k-means, and compares pairs by cosine similarity.
What NVIDIA says (2)
“uses embeddings to identify and remove “semantic duplicates” - data pairs that are semantically similar but not exactly identical.”
“The embeddings are clustered into k clusters using k-means clustering.”
Cosine similarity measures how closely two vectors point the same way. Pairs above the chosen threshold are treated as semantic duplicates.
What NVIDIA says (1)
“Data pairs with cosine similarity above a threshold are considered semantic duplicates.”
A heuristic is a simple rule of thumb. Heuristic filters compute simple statistics and drop documents that fail them. This also applies to the captions in image-text data.
What NVIDIA says (1)
“There are heuristics that measure quality by gathering simple statistics like how many punctutation marks a document has, how long is the document, and how repetitive is the document.”
PII means personal identifiable information, such as names, emails and phone numbers. Curator's tool finds and removes it at scale with Dask.
What NVIDIA says (1)
“The purpose of the personal identifiable information (PII) de-identification tool is to help scrub sensitive data out of datasets.”
An image embedding is a vector that summarizes an image. Quality scoring and duplicate checks then run on these vectors, which is far cheaper than on pixels.
What NVIDIA says (1)
“Image embeddings are the backbone to many data curation operations in NeMo Curator.”
Key terms: NeMo Curator Aesthetic classifier NSFW classifier Semantic deduplication Personal identifiable information Embedding
Try it: Deduplication lab
1.5 Testing models for accuracy
FID-CLIP, retrieval metrics, perplexity, BLEU, LLM-as-a-judge, cross-validation and robust error metrics.
Key points
FID (Fréchet Inception Distance) measures how close generated images are to real ones. A CLIP score measures how well images match their prompts. NVIDIA found both UNets similar on FID-CLIP, with slightly better visual quality from the Regular UNet. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”
Retrieval metrics check whether the right items come back. Recall@K counts how many of the relevant items appear in the top K. Precision@K counts how many of the top K are relevant. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“Assesses the proportion of relevant documents that are successfully retrieved in a grouping of K retrieved documents.”
LLM-as-a-judge uses one LLM to grade another model's output. It helps where simple metrics fall short. It can inherit bias from the judge model. LLM means large language model.
What NVIDIA says (2)
“This method excels for tasks where automated metrics fall short, such as assessing coherence and creativity.”
“it’s important to note that LLM-as-a-judge evaluations may introduce biases inherent in the LLM evaluator training data.”
Cross-validation trains and validates several times on different splits. This gives a steadier estimate of accuracy than one split.
What NVIDIA says (1)
“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”
MAE averages the absolute errors. Because it does not square them, large errors do not dominate.
What NVIDIA says (1)
“All errors are treated equally, so the metric is robust to outliers.”
Precision@K looks at what came back in the top K. Recall@K looks at how many relevant items were found. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“Measures the proportion of retrieved documents that are relevant in a grouping of K retrieved documents.”
Perplexity measures how uncertain a model is when predicting the next words. Lower uncertainty means better prediction.
What NVIDIA says (1)
“Quantifies uncertainty in predicting sequences of words, with lower values indicating better predictive performance.”
BLEU (Bilingual Evaluation Understudy) compares model output with reference translations. MAE means mean absolute error.
What NVIDIA says (2)
“Evaluates machine translation quality by comparing model outputs to reference translations.”
“The BLEU score ranges from 0 (no match, that is, low quality) to 1 (a perfect match, that is, high quality).”
Key terms: Recall@K Precision@K FID-CLIP evaluation LLM-as-a-judge Cross-validation Mean absolute error Perplexity
Try it: Metrics lab