Experimentation

25% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI

1.1 Developing and testing multimodal models

Official objective: “Assist in developing and testing multimodal AI models.”

How NeVA, CLIP, Stable Diffusion, Video NeVA and SpeechLLMs are put together, and what to check when you build one.

Key points

  1. A multimodal model works with more than one kind of data, such as images and text. NeVA (NeMo Vision and Language Assistant) joins a large language model (LLM) to a vision encoder. A vision encoder is a network that turns an image into feature vectors.

    What NVIDIA says (1)

    “It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”

    — NeMo Framework 24.09: NeVA

  2. The vision encoder in NeVA is the pretrained CLIP ViT-L/14. A projection matrix is a learned layer that changes vectors from one space into another. NeVA uses it to blend visual features with the language embeddings. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (2)

    “NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”

    — NeMo Framework 24.09: NeVA

    “The encoder retrieves visual features from images and intertwines them with language embeddings using a modifiable projection matrix.”

    — NeMo Framework 24.09: NeVA

  3. CLIP (Contrastive Language-Image Pre-training) learns from image-caption pairs. Contrastive training pulls matching pairs together and pushes mismatched pairs apart. The result is a shared space where images and text can be compared.

    What NVIDIA says (2)

    “The essence of CLIP is to train both an image encoder and a text encoder from scratch.”

    — NeMo Framework 24.09: CLIP

    “maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

    — NeMo Framework 24.09: CLIP

  4. Stable Diffusion generates images from text. The U-Net predicts noise. The variational autoencoder (VAE) compresses images into a smaller latent space. The CLIP text encoder turns the prompt into embeddings that guide the U-Net. CLIP means Contrastive Language-Image Pre-training. GAN means generative adversarial network.

    What NVIDIA says (1)

    “Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”

    — NeMo Framework 24.09: Stable Diffusion

  5. A modality is a type of data, such as text, image, audio or video. Video NeVA treats a video as a series of frames. Its config sets how many frames to take, for example num_frames.

    What NVIDIA says (1)

    “Video NeVa adds support for video modality in NeVa by representing video as multiple image frames.”

    — NeMo Framework 24.09: Video NeVA

  6. A SpeechLLM is a large language model that also accepts audio. An audio encoder turns sound into audio embeddings. The modality adapter maps those into the LLM's embedding space so the LLM can use them with the text prompt.

    What NVIDIA says (1)

    “A modality adapter that processes the audio embeddings and produces a sequence of embeddings in the same latent space as the token embeddings of a pretrained LLM.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

  7. Concatenation places the speech embeddings next to the text embeddings in time. Cross-attention lets one sequence (text) look up information in another (speech). NeMo adds the cross-attention module only before the LLM to keep cost down. LLM means large language model.

    What NVIDIA says (3)

    “One way to incorporate speech into an LLM is to concatenate speech features with the token embeddings of the input text prompt before feeding them into the LLM.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

    “The Speech-Augmented Language Model (SALM) follows this approach.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

    “Another approach is to use a cross-attention mechanism, where text embeddings attend to speech embeddings to extract task-specific information.”

    — NeMo Framework 24.09: Speech-agnostic Multimodal LLMs (SpeechLLM)

  8. The projection maps image features into the language model's embedding space. A multilayer perceptron (MLP) is a small stack of fully connected layers. A two-layer MLP is more expressive than one linear layer.

    What NVIDIA says (1)

    “Transitioning from a linear to a dual-layer MLP projection markedly bolsters LLaVA-1.5’s multimodal faculties”

    — NeMo Framework 24.09: NeVA

Key terms: Multimodal model Modality NeVA Vision encoder Projection layer CLIP Stable Diffusion SpeechLLM Modality adapter Cross-attention

Try it: CLIP lab

Practice 1.1 (8 questions)

1.2 Managing and preprocessing multimodal data

Official objective: “Manage and preprocess data from various sources.”

WebDataset shards, the two NeVA data phases, precomputed encodings and GPU data loading with DALI.

Key points

  1. Pretraining aligns image features with the language model on many caption pairs. Fine-tuning teaches the model to follow instructions about images. Each phase needs a different dataset.

    What NVIDIA says (1)

    “The NeVA model training encompasses two phases: pretraining and fine-tuning. Each phase mandates a unique dataset.”

    — NeMo Framework 24.09: Multimodal Language Model Datasets

  2. WebDataset is a format that packs samples into tar archives called shards. Equal-size shards stream well during training.

    What NVIDIA says (1)

    “The pipeline processes the dataset into the WebDataset format, consisting of tar files of equal sizes for efficient training.”

    — NeMo Framework 24.09: Text to Image Datasets

  3. Images can fail to download because of network errors or removed files. That leaves uneven tar files. The optional reorganize_tar stage repacks them so each holds the same number of image-text pairs.

    What NVIDIA says (2)

    “some images may fail to download, resulting in uneven tar files with varying number of examples each.”

    — NeMo Framework 24.09: Text to Image Datasets

    “you need to re-organize the contents of the tar files so that each one contains an equal number of image-text pairs.”

    — NeMo Framework 24.09: Text to Image Datasets

  4. A frozen encoder is one whose weights do not change in training. Its output for a given input never changes, so you can compute it once. NeMo calls this precaching the encodings. VAE means variational autoencoder.

    What NVIDIA says (2)

    “Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”

    — NeMo Framework 24.09: Stable Diffusion

    “Precaching these encodings can significantly enhance training throughput.”

    — NeMo Framework 24.09: Text to Image Datasets

  5. DALI (Data Loading Library) loads and preprocesses image, video and audio data. It offloads that work from the CPU to the GPU. DALI means the NVIDIA Data Loading Library. SDK means software development kit. LLM means large language model.

    What NVIDIA says (2)

    “The NVIDIA Data Loading Library (DALI) is a GPU-accelerated library for data loading and pre-processing to accelerate deep learning applications.”

    — NVIDIA DALI User Guide

    “DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the GPU.”

    — NVIDIA DALI User Guide

  6. DALI is built for the media types used in multimodal training. It provides optimized building blocks for images, video and audio. DALI means the NVIDIA Data Loading Library.

    What NVIDIA says (1)

    “It provides a collection of highly optimized building blocks for loading and processing image, video and audio data.”

    — NVIDIA DALI User Guide

  7. A dataset license sets how the data may be used. NVIDIA's docs say each user must review the content and license before use.

    What NVIDIA says (1)

    “It is the responsibility of each user to check the content of the dataset, review the applicable licenses, and determine if it is suitable for their intended use.”

    — NeMo Framework 24.09: Text to Image Datasets

Key terms: WebDataset NVIDIA DALI

Practice 1.2 (7 questions)

1.3 Explainability with multimodal models

Official objective: “Use multimodal models to improve explainability.”

Attention, LIME and SHAP, grounding in RAG, model cards and proxy models.

Key points

  1. Explainable AI (XAI) is a set of tools and techniques that help people understand why a model decided something. For images, audio and text, attention can be visualized to show which parts of the input mattered.

    What NVIDIA says (2)

    “is a set of tools and techniques used by organizations to help people better understand why a model makes certain decisions and how it works.”

    — What Is Explainable AI (XAI)?

    “For some data — images, audio and text — similar results can be visualized through the use of”

    — What Is Explainable AI (XAI)?

  2. LIME and SHAP assign credit to each input feature for one prediction. NVIDIA notes they give literal mathematical answers that can be shown to many audiences.

    What NVIDIA says (1)

    “Techniques with names like LIME and SHAP offer very literal mathematical answers to this question”

    — What Is Explainable AI (XAI)?

  3. Retrieval-augmented generation (RAG) fetches relevant data and gives it to the model as context. Grounding means converting other modalities into one primary modality, here text. The text descriptions also make the retrieved evidence readable by people.

    What NVIDIA says (2)

    “The key benefit here is that the metadata generated from the information-rich image is extremely helpful in answering objective questions.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

    “The key disadvantages are preprocessing costs and losing some nuance from the image.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

  4. RAG retrieves documents and passes them to the model. The model can then cite them, so users can verify each claim. A multimodal RAG app can cite retrieved images and charts the same way.

    What NVIDIA says (1)

    “Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”

    — What Is Retrieval-Augmented Generation aka RAG

  5. A model card is a short document that describes a model, its uses and its limits. Model Card++ adds four subsections on trust topics.

    What NVIDIA says (1)

    “Four subsections detailing model-specific information concerning Bias, Explainability, Privacy, and Safety and Security.”

    — Enhancing AI Transparency and Ethical Considerations with Model Card++

  6. Proxy modeling uses a simple model to approximate a complex one. It gives a sense of the whole model, but it is only an approximation.

    What NVIDIA says (1)

    “simpler, more easily comprehended models like decision trees can be used to approximately describe the more detailed AI model.”

    — What Is Explainable AI (XAI)?

  7. Pedigree is the history of how a model was made. Knowing it helps people judge when the model's outputs make sense, including for multimodal models trained on images and text.

    What NVIDIA says (1)

    “Explaining the pedigree of the model: How was the model trained? What data was used? How was the impact of any bias in the training data measured and mitigated?”

    — What Is Explainable AI (XAI)?

Key terms: Attention map Explainable AI Model card Retrieval-augmented generation

Practice 1.3 (7 questions)

1.4 Testing multimodal data quality

Official objective: “Test data quality and consistency in a multimodal setting.”

NeMo Curator aesthetic and NSFW scores, semantic deduplication, heuristic filters and PII removal.

Key points

  1. NeMo Curator is NVIDIA's data curation library. Its aesthetic classifier estimates the subjective quality of an image. A higher score means a more pleasing image. PII means personal identifiable information. NSFW means not safe for work.

    What NVIDIA says (2)

    “Aesthetic classifiers can be used to assess the subjective quality of an image.”

    — NeMo Curator (NeMo Framework 24.09): Aesthetic Classifier

    “outputs a score from 0-10 where 10 is aesthetically pleasing.”

    — NeMo Curator (NeMo Framework 24.09): Aesthetic Classifier

  2. An embedding is a vector that represents an input. The aesthetic classifier is a small linear model on top of CLIP image embeddings. CLIP means Contrastive Language-Image Pre-training. PII means personal identifiable information.

    What NVIDIA says (1)

    “a linear classifier that takes OpenAI CLIP ViT-L/14 image embeddings as input.”

    — NeMo Curator (NeMo Framework 24.09): Aesthetic Classifier

  3. NSFW means not safe for work. Removing unsafe content is a common step in generative AI data pipelines. PII means personal identifiable information.

    What NVIDIA says (2)

    “a value between 0 and 1 where 1 means the content is NSFW.”

    — NeMo Curator (NeMo Framework 24.09): NSFW Classifier

    “Removing unsafe content is common in most data processing pipelines to prevent your generative AI model from learning to produce unsafe material.”

    — NeMo Curator (NeMo Framework 24.09): NSFW Classifier

  4. Semantic duplicates are pairs that mean almost the same thing without being identical. Semantic deduplication embeds each sample, clusters the embeddings with k-means, and compares pairs by cosine similarity.

    What NVIDIA says (2)

    “uses embeddings to identify and remove “semantic duplicates” - data pairs that are semantically similar but not exactly identical.”

    — NeMo Curator (NeMo Framework 24.09): Semantic Deduplication

    “The embeddings are clustered into k clusters using k-means clustering.”

    — NeMo Curator (NeMo Framework 24.09): Semantic Deduplication

  5. Cosine similarity measures how closely two vectors point the same way. Pairs above the chosen threshold are treated as semantic duplicates.

    What NVIDIA says (1)

    “Data pairs with cosine similarity above a threshold are considered semantic duplicates.”

    — NeMo Curator (NeMo Framework 24.09): Semantic Deduplication

  6. A heuristic is a simple rule of thumb. Heuristic filters compute simple statistics and drop documents that fail them. This also applies to the captions in image-text data.

    What NVIDIA says (1)

    “There are heuristics that measure quality by gathering simple statistics like how many punctutation marks a document has, how long is the document, and how repetitive is the document.”

    — NeMo Curator (NeMo Framework 24.09): Classifier and Heuristic Quality Filtering

  7. PII means personal identifiable information, such as names, emails and phone numbers. Curator's tool finds and removes it at scale with Dask.

    What NVIDIA says (1)

    “The purpose of the personal identifiable information (PII) de-identification tool is to help scrub sensitive data out of datasets.”

    — NeMo Curator (NeMo Framework 24.09): PII Identification and Removal

  8. An image embedding is a vector that summarizes an image. Quality scoring and duplicate checks then run on these vectors, which is far cheaper than on pixels.

    What NVIDIA says (1)

    “Image embeddings are the backbone to many data curation operations in NeMo Curator.”

    — NeMo Curator (NeMo Framework 24.09): Image Curation

Key terms: NeMo Curator Aesthetic classifier NSFW classifier Semantic deduplication Personal identifiable information Embedding

Try it: Deduplication lab

Practice 1.4 (8 questions)

1.5 Testing models for accuracy

Official objective: “Test AI models to ensure their accuracy and effectiveness.”

FID-CLIP, retrieval metrics, perplexity, BLEU, LLM-as-a-judge, cross-validation and robust error metrics.

Key points

  1. FID (Fréchet Inception Distance) measures how close generated images are to real ones. A CLIP score measures how well images match their prompts. NVIDIA found both UNets similar on FID-CLIP, with slightly better visual quality from the Regular UNet. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”

    — NeMo Framework 24.09: Imagen

  2. Retrieval metrics check whether the right items come back. Recall@K counts how many of the relevant items appear in the top K. Precision@K counts how many of the top K are relevant. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “Assesses the proportion of relevant documents that are successfully retrieved in a grouping of K retrieved documents.”

    — Mastering LLM Techniques: Evaluation

  3. LLM-as-a-judge uses one LLM to grade another model's output. It helps where simple metrics fall short. It can inherit bias from the judge model. LLM means large language model.

    What NVIDIA says (2)

    “This method excels for tasks where automated metrics fall short, such as assessing coherence and creativity.”

    — Mastering LLM Techniques: Evaluation

    “it’s important to note that LLM-as-a-judge evaluations may introduce biases inherent in the LLM evaluator training data.”

    — Mastering LLM Techniques: Evaluation

  4. Cross-validation trains and validates several times on different splits. This gives a steadier estimate of accuracy than one split.

    What NVIDIA says (1)

    “Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”

    — What is scikit-learn?

  5. MAE averages the absolute errors. Because it does not square them, large errors do not dominate.

    What NVIDIA says (1)

    “All errors are treated equally, so the metric is robust to outliers.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  6. Precision@K looks at what came back in the top K. Recall@K looks at how many relevant items were found. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “Measures the proportion of retrieved documents that are relevant in a grouping of K retrieved documents.”

    — Mastering LLM Techniques: Evaluation

  7. Perplexity measures how uncertain a model is when predicting the next words. Lower uncertainty means better prediction.

    What NVIDIA says (1)

    “Quantifies uncertainty in predicting sequences of words, with lower values indicating better predictive performance.”

    — Mastering LLM Techniques: Evaluation

  8. BLEU (Bilingual Evaluation Understudy) compares model output with reference translations. MAE means mean absolute error.

    What NVIDIA says (2)

    “Evaluates machine translation quality by comparing model outputs to reference translations.”

    — Mastering LLM Techniques: Evaluation

    “The BLEU score ranges from 0 (no match, that is, low quality) to 1 (a perfect match, that is, high quality).”

    — Mastering LLM Techniques: Evaluation

Key terms: Recall@K Precision@K FID-CLIP evaluation LLM-as-a-judge Cross-validation Mean absolute error Perplexity

Try it: Metrics lab

Practice 1.5 (8 questions)