Multimodal Data

15% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI

3.1 Scalability, performance and reliability

Official objective: “Assist in the deployment and evaluations of model scalability, performance, and reliability under the supervision of a senior team member.”

Dynamic batching, load tests versus benchmarks, and readiness probes.

Key points

  1. Triton Inference Server is NVIDIA's model serving software. Dynamic batching combines separate requests into one batch on the server. Bigger batches use the GPU better.

    What NVIDIA says (1)

    “Dynamic batching is a feature of Triton that allows inference requests to be combined by the server, so that a batch is created dynamically. Creating a batch of requests typically results in increased throughput.”

    — NVIDIA Triton Inference Server: Batchers

  2. By default the batcher does not wait for more requests. You can set a queue delay if you prefer bigger batches over latency.

    What NVIDIA says (1)

    “By default the dynamic batcher will create batches as large as possible up to the maximum batch size and will not delay when forming batches.”

    — NVIDIA Triton Inference Server: Batchers

  3. Scalability is how well a service handles more traffic. Load tests probe scale. Benchmarks such as GenAI-Perf measure raw model performance. NVIDIA recommends both.

    What NVIDIA says (1)

    “Load testing focuses on simulating a large number of concurrent requests to a model to assess its ability to handle real-world traffic at scale.”

    — LLM Inference Benchmarking: Fundamental Concepts

  4. A readiness probe checks whether a service can take traffic. NIM for VLMs returns 200 on /v1/health/ready once the model is loaded. VLM means vision language model. VLMs means vision language models. NIM means NVIDIA Inference Microservice.

    What NVIDIA says (1)

    “GET /v1/health/ready Readiness probe. Returns 200 when the model is loaded and inference is available.”

    — NVIDIA NIM for Vision Language Models: API Reference

Key terms: Triton Inference Server Dynamic batching NVIDIA NIM Readiness probe

Practice 3.1 (4 questions)

3.2 RAG, chatbots and summarizers

Official objective: “Build LLM use cases such as retrieval-augmented generation (RAG), chatbots, and summarizers.”

Multimodal RAG designs, the offline and online parts of RAG, and hallucination.

Key points

  1. CLIP can encode both text and images into the same space. The rest of the text RAG setup can stay much the same. A multimodal LLM (MLLM) then answers using retrieved images. CLIP means Contrastive Language-Image Pre-training. RAG means retrieval-augmented generation. LLM means large language model. MLLM means multimodal large language model.

    What NVIDIA says (2)

    “In the case of images and text, you can use a model like CLIP to encode both text and images in the same vector space.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

    “For the generation pass, you then replace the large language model (LLM) with a multimodal LLM (MLLM) for all question and answering.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

  2. Each approach trades simplicity against fidelity. Separate stores need a rank-rerank step to merge results. RAG means retrieval-augmented generation. OCR means optical character recognition.

    What NVIDIA says (1)

    “Embed all modalities into the same vector space Ground all modalities into one primary modality Have separate stores for different modalities”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

  3. Ingestion loads, splits and embeds documents into a vector database. Retrieval and generation happen online when a query arrives. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

  4. Hallucination is a plausible but incorrect answer. RAG supplies real sources as context, which lowers that risk. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”

    — What Is Retrieval-Augmented Generation aka RAG

Key terms: Modality Retrieval-augmented generation Hallucination

Try it: RAG lab

Practice 3.2 (4 questions)

3.3 Python natural-language packages

Official objective: “Familiarity with the capabilities of Python natural language packages (spaCy, NumPy, vector databases, etc.).”

Vector databases, similarity search, spaCy and NumPy.

Key points

  1. An embedding is a list of numbers that captures meaning. A vector database stores embeddings and finds the ones closest to a query.

    What NVIDIA says (1)

    “A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”

    — What is a Vector Database and How Does it Work?

  2. Similarity search is also called vector or semantic search. It ranks stored vectors by closeness to the query vector.

    What NVIDIA says (1)

    “refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”

    — What is a Vector Database and How Does it Work?

  3. spaCy is a popular Python library for natural-language processing, such as named-entity recognition (NER). Moving data back to the CPU slows a GPU pipeline.

    What NVIDIA says (1)

    “relied on spaCy for NER but, spaCy currently needs your inputs on CPU”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

  4. NumPy is the core Python library for arrays and linear algebra. Many data and NLP libraries build on it. NLP means natural-language processing.

    What NVIDIA says (1)

    “scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”

    — What is scikit-learn?

Key terms: Vector database Embedding spaCy NumPy

Try it: RAG lab

Practice 3.3 (4 questions)

3.4 Identifying data, hardware and software components

Official objective: “Identify system data, hardware, or software components required to meet user needs.”

Text encoders, latent spaces, speech SDKs and number formats that fit GPU memory.

Key points

  1. A text encoder turns a prompt into embeddings. Imagen uses a large language model encoder, typically T5, instead of CLIP. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Imagen employs a text encoder, typically T5, to encode textual features.”

    — NeMo Framework 24.09: Imagen

  2. A latent space is a compressed representation of the data. Diffusing in that smaller space is cheaper than in pixel space. VAE means variational autoencoder.

    What NVIDIA says (2)

    “an input image is condensed from 512x512x3 dimensions to 64x64x4.”

    — NeMo Framework 24.09: Stable Diffusion

    “This compression results in decreased memory and computational requirements when compared to pixel-space diffusion models.”

    — NeMo Framework 24.09: Stable Diffusion

  3. Automatic speech recognition (ASR) turns audio into text. Text-to-speech (TTS) turns text into audio. Riva provides both on GPUs. DALI means the NVIDIA Data Loading Library. SDK means software development kit.

    What NVIDIA says (2)

    “Automatic Speech Recognition (ASR) takes an audio stream or audio buffer as input and returns one or more text transcripts”

    — NVIDIA Riva: ASR Overview

    “based on a two-stage pipeline.”

    — NVIDIA Riva: TTS Overview

  4. FP16 is half-precision floating point. Half the bits means roughly half the memory for the same values. FP16 means 16-bit floating point.

    What NVIDIA says (1)

    “Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32). Lowering the required memory enables training of larger models or training with larger mini-batches.”

    — Train With Mixed Precision

Key terms: Variational autoencoder Latent space Text encoder Imagen Mixed precision NVIDIA Riva Automatic speech recognition Text-to-speech T5

Practice 3.4 (4 questions)

3.5 Monitoring data collection and experiments

Official objective: “Monitor the functioning of data collection, experiments, and other software processes.”

Experiment loggers, resuming runs, MLOps tracking and data flywheels.

Key points

  1. Experiment tracking records metrics, settings and checkpoints for each run. NeMo's exp_manager sets this up through PyTorch Lightning.

    What NVIDIA says (1)

    “The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”

    — NeMo Framework 24.09: Experiment Manager

  2. A checkpoint is a saved copy of model and optimizer state. With resume_if_exists, the run picks up from the latest checkpoint.

    What NVIDIA says (1)

    “resume training if checkpoints already exist resume_if_exists : True”

    — NeMo Framework 24.09: Experiment Manager

  3. MLOps (machine learning operations) is a set of practices for running AI reliably. Tracking records which data and settings produced each model.

    What NVIDIA says (1)

    “AI models require careful tracking through cycles of experiments, tuning and retraining.”

    — What is MLOps?

  4. Monitoring what users do with a model creates new data. Feeding that data back improves the model over time.

    What NVIDIA says (1)

    “An AI data flywheel is a self-improving loop where data collected from AI interactions or processes is used to continuously refine AI models”

    — Data flywheel: What it is and how it works

Key terms: Experiment Manager Checkpoint MLOps Data flywheel

Practice 3.5 (4 questions)

3.6 Python packages for traditional ML

Official objective: “Use Python packages (spaCy, NumPy, Keras, etc.) to implement specific traditional machine learning analyses.”

XGBoost, pandas and scikit-learn estimators.

Key points

  1. Gradient boosting builds many small decision trees, each fixing the errors of the last. XGBoost is a popular library for it on tabular data.

    What NVIDIA says (1)

    “is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”

    — What Is XGBoost and Why Does It Matter?

  2. pandas provides DataFrames, which are tables with labeled rows and columns. It is a common first step before traditional ML. ML means machine learning.

    What NVIDIA says (1)

    “pandas is the most popular software library for data manipulation and data analysis for the Python programming languages.”

    — What Is Pandas and Why Does it Matter?

  3. Every scikit-learn model, such as a random forest, is an estimator with fit and predict methods.

    What NVIDIA says (1)

    “An estimator is the core machine learning algorithm that fits the training data to produce a model.”

    — What is scikit-learn?

Key terms: scikit-learn Estimator XGBoost pandas

Practice 3.6 (3 questions)

3.7 Writing software components and scripts

Official objective: “Write software components or scripts under the supervision of a senior team member.”

Calling a NIM VLM through its OpenAI-compatible API and starting Triton with a model repository.

Key points

  1. OpenAI-compatible means existing OpenAI client code can call it by changing the base URL. Chat completions take a message history and return the model's reply. VLM means vision language model. NIM means NVIDIA Inference Microservice.

    What NVIDIA says (2)

    “NIM VLM exposes an OpenAI-compatible inference API backed by vLLM”

    — NVIDIA NIM for Vision Language Models: API Reference

    “POST /v1/chat/completions Multi-turn chat completions with message history.”

    — NVIDIA NIM for Vision Language Models: API Reference

  2. A model repository is a folder with each model and its configuration. Triton loads models from it at startup.

    What NVIDIA says (1)

    “tritonserver --model-repository=/models”

    — Quickstart — NVIDIA Triton Inference Server

Key terms: Triton Inference Server NVIDIA NIM

Practice 3.7 (2 questions)