Multimodal Data
15% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI
3.1 Scalability, performance and reliability
Dynamic batching, load tests versus benchmarks, and readiness probes.
Key points
Triton Inference Server is NVIDIA's model serving software. Dynamic batching combines separate requests into one batch on the server. Bigger batches use the GPU better.
What NVIDIA says (1)
“Dynamic batching is a feature of Triton that allows inference requests to be combined by the server, so that a batch is created dynamically. Creating a batch of requests typically results in increased throughput.”
By default the batcher does not wait for more requests. You can set a queue delay if you prefer bigger batches over latency.
What NVIDIA says (1)
“By default the dynamic batcher will create batches as large as possible up to the maximum batch size and will not delay when forming batches.”
Scalability is how well a service handles more traffic. Load tests probe scale. Benchmarks such as GenAI-Perf measure raw model performance. NVIDIA recommends both.
What NVIDIA says (1)
“Load testing focuses on simulating a large number of concurrent requests to a model to assess its ability to handle real-world traffic at scale.”
A readiness probe checks whether a service can take traffic. NIM for VLMs returns 200 on /v1/health/ready once the model is loaded. VLM means vision language model. VLMs means vision language models. NIM means NVIDIA Inference Microservice.
What NVIDIA says (1)
“GET /v1/health/ready Readiness probe. Returns 200 when the model is loaded and inference is available.”
Key terms: Triton Inference Server Dynamic batching NVIDIA NIM Readiness probe
3.2 RAG, chatbots and summarizers
Multimodal RAG designs, the offline and online parts of RAG, and hallucination.
Key points
CLIP can encode both text and images into the same space. The rest of the text RAG setup can stay much the same. A multimodal LLM (MLLM) then answers using retrieved images. CLIP means Contrastive Language-Image Pre-training. RAG means retrieval-augmented generation. LLM means large language model. MLLM means multimodal large language model.
What NVIDIA says (2)
“In the case of images and text, you can use a model like CLIP to encode both text and images in the same vector space.”
“For the generation pass, you then replace the large language model (LLM) with a multimodal LLM (MLLM) for all question and answering.”
Each approach trades simplicity against fidelity. Separate stores need a rank-rerank step to merge results. RAG means retrieval-augmented generation. OCR means optical character recognition.
What NVIDIA says (1)
“Embed all modalities into the same vector space Ground all modalities into one primary modality Have separate stores for different modalities”
Ingestion loads, splits and embeds documents into a vector database. Retrieval and generation happen online when a query arrives. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
Hallucination is a plausible but incorrect answer. RAG supplies real sources as context, which lowers that risk. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”
Key terms: Modality Retrieval-augmented generation Hallucination
Try it: RAG lab
3.3 Python natural-language packages
Vector databases, similarity search, spaCy and NumPy.
Key points
An embedding is a list of numbers that captures meaning. A vector database stores embeddings and finds the ones closest to a query.
What NVIDIA says (1)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
Similarity search is also called vector or semantic search. It ranks stored vectors by closeness to the query vector.
What NVIDIA says (1)
“refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
spaCy is a popular Python library for natural-language processing, such as named-entity recognition (NER). Moving data back to the CPU slows a GPU pipeline.
What NVIDIA says (1)
“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”
NumPy is the core Python library for arrays and linear algebra. Many data and NLP libraries build on it. NLP means natural-language processing.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
Key terms: Vector database Embedding spaCy NumPy
Try it: RAG lab
3.4 Identifying data, hardware and software components
Text encoders, latent spaces, speech SDKs and number formats that fit GPU memory.
Key points
A text encoder turns a prompt into embeddings. Imagen uses a large language model encoder, typically T5, instead of CLIP. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Imagen employs a text encoder, typically T5, to encode textual features.”
A latent space is a compressed representation of the data. Diffusing in that smaller space is cheaper than in pixel space. VAE means variational autoencoder.
What NVIDIA says (2)
“an input image is condensed from 512x512x3 dimensions to 64x64x4.”
“This compression results in decreased memory and computational requirements when compared to pixel-space diffusion models.”
Automatic speech recognition (ASR) turns audio into text. Text-to-speech (TTS) turns text into audio. Riva provides both on GPUs. DALI means the NVIDIA Data Loading Library. SDK means software development kit.
What NVIDIA says (2)
“Automatic Speech Recognition (ASR) takes an audio stream or audio buffer as input and returns one or more text transcripts”
“based on a two-stage pipeline.”
FP16 is half-precision floating point. Half the bits means roughly half the memory for the same values. FP16 means 16-bit floating point.
What NVIDIA says (1)
“Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32). Lowering the required memory enables training of larger models or training with larger mini-batches.”
Key terms: Variational autoencoder Latent space Text encoder Imagen Mixed precision NVIDIA Riva Automatic speech recognition Text-to-speech T5
3.5 Monitoring data collection and experiments
Experiment loggers, resuming runs, MLOps tracking and data flywheels.
Key points
Experiment tracking records metrics, settings and checkpoints for each run. NeMo's exp_manager sets this up through PyTorch Lightning.
What NVIDIA says (1)
“The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”
A checkpoint is a saved copy of model and optimizer state. With resume_if_exists, the run picks up from the latest checkpoint.
What NVIDIA says (1)
“resume training if checkpoints already exist resume_if_exists : True”
MLOps (machine learning operations) is a set of practices for running AI reliably. Tracking records which data and settings produced each model.
What NVIDIA says (1)
“AI models require careful tracking through cycles of experiments, tuning and retraining.”
Monitoring what users do with a model creates new data. Feeding that data back improves the model over time.
What NVIDIA says (1)
“An AI data flywheel is a self-improving loop where data collected from AI interactions or processes is used to continuously refine AI models”
Key terms: Experiment Manager Checkpoint MLOps Data flywheel
3.6 Python packages for traditional ML
XGBoost, pandas and scikit-learn estimators.
Key points
Gradient boosting builds many small decision trees, each fixing the errors of the last. XGBoost is a popular library for it on tabular data.
What NVIDIA says (1)
“is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”
pandas provides DataFrames, which are tables with labeled rows and columns. It is a common first step before traditional ML. ML means machine learning.
What NVIDIA says (1)
“pandas is the most popular software library for data manipulation and data analysis for the Python programming languages.”
Every scikit-learn model, such as a random forest, is an estimator with fit and predict methods.
What NVIDIA says (1)
“An estimator is the core machine learning algorithm that fits the training data to produce a model.”
Key terms: scikit-learn Estimator XGBoost pandas
3.7 Writing software components and scripts
Calling a NIM VLM through its OpenAI-compatible API and starting Triton with a model repository.
Key points
OpenAI-compatible means existing OpenAI client code can call it by changing the base URL. Chat completions take a message history and return the model's reply. VLM means vision language model. NIM means NVIDIA Inference Microservice.
What NVIDIA says (2)
“NIM VLM exposes an OpenAI-compatible inference API backed by vLLM”
“POST /v1/chat/completions Multi-turn chat completions with message history.”
A model repository is a folder with each model and its configuration. Triton loads models from it at startup.
What NVIDIA says (1)
“tritonserver --model-repository=/models”
Key terms: Triton Inference Server NVIDIA NIM