NCA-GENM glossary
The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.
A
- Activation recomputation
Storing only some activations and recomputing the rest in the backward pass to save memory.
What NVIDIA says (2)
“Checkpointing a few activations and recomputing the rest is a common technique to reduce device memory usage.”
“it increases the per-transformer layer computation cost by 30%”
- Aesthetic classifier
A NeMo Curator model that scores image quality from 0 to 10 from CLIP image embeddings.
What NVIDIA says (2)
“outputs a score from 0-10 where 10 is aesthetically pleasing.”
“a linear classifier that takes OpenAI CLIP ViT-L/14 image embeddings as input.”
- Attention map
A picture of attention weights showing which parts of the input a model focused on.
What NVIDIA says (2)
“Our main insight is that the key pathway of the cross-attention module in the diffusion model (the K matrix) controls the layout of the attention maps.”
“Attention units follow these tags, calculating a kind of algebraic map of how each element relates to the others.”
- Automatic mixed precision
A framework feature that picks 16-bit or 32-bit math per operation automatically.
What NVIDIA says (1)
“Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”
- Automatic speech recognition
Turning audio into text.
What NVIDIA says (1)
“Automatic Speech Recognition (ASR) takes an audio stream or audio buffer as input and returns one or more text transcripts”
B
- bfloat16
A 16-bit number format with FP32's range, used to train faster with less memory.
What NVIDIA says (1)
“Training is conducted in Bfloat16 precision, which offers a balance between the higher precision of FP32 and the memory savings and speed of FP16.”
C
- Checkpoint
A saved copy of training state so a run can resume.
What NVIDIA says (1)
“resume training if checkpoints already exist resume_if_exists : True”
- CLIP
A model with an image encoder and a text encoder trained together so matching images and captions land close in one vector space.
What NVIDIA says (2)
“The essence of CLIP is to train both an image encoder and a text encoder from scratch.”
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
- Clustering
Grouping similar items without labels; a common unsupervised task.
What NVIDIA says (1)
“Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”
- Confidential computing
Protecting code, models and data while in use, with hardware trusted execution environments.
What NVIDIA says (1)
“Confidential computing secures AI workloads by encrypting application code, models, and data in use—not just at rest or in transit.”
- Confounding factor
A hidden variable that affects both inputs and outcomes and can create a false link.
What NVIDIA says (1)
“Medical institutions have had to rely on their own data sources, which can be biased by, for example, patient demographics, the instruments used or clinical specializations.”
- Consent
A person's agreement to a specific use of their data.
What NVIDIA says (2)
“Developers of AI models that rely on data such as a person’s image, voice, artistic work or health records”
“should evaluate whether individuals have provided appropriate consent for their personal information to be used in this way.”
- Contrastive loss
A loss that rewards high similarity for matching pairs and low similarity for mismatched pairs.
What NVIDIA says (1)
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
- ControlNet
A side network that adds conditions such as edge maps to a diffusion model through a locked copy and a trainable copy of its blocks.
What NVIDIA says (1)
“It copies the weights of neural network blocks into a “locked” copy and a “trainable” copy.”
- Correlation
A measure of how two variables move together; it does not prove cause.
What NVIDIA says (2)
“For any use case, be aware of the variables that are highly correlated, as they could skew results.”
“Take this into account if you are using any future algorithm that assumes the variables are independent, such as linear regression .”
- Cross-attention
Attention in which one sequence, such as text or image features, looks up information in another sequence.
What NVIDIA says (2)
“Another approach is to use a cross-attention mechanism, where text embeddings attend to speech embeddings to extract task-specific information.”
“Our main insight is that the key pathway of the cross-attention module in the diffusion model (the K matrix) controls the layout of the attention maps.”
- Cross-filtering
Linked charts where selecting data in one filters all the others.
What NVIDIA says (1)
“cuxfilter enables GPU accelerated cross-filtering dashboards from notebooks, in just a few lines of Python code.”
- Cross-validation
Training and validating several times on different splits of the data to estimate accuracy fairly.
What NVIDIA says (2)
“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”
“For effective model selection, scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”
D
- Data flywheel
A loop where data from AI use improves the model, which then produces better data.
What NVIDIA says (1)
“An AI data flywheel is a self-improving loop where data collected from AI interactions or processes is used to continuously refine AI models”
- Data parallelism
Copying a model to several GPUs and splitting each batch between them.
What NVIDIA says (1)
“Data Parallelism (DP) replicates the model across multiple GPUs. Data batches are evenly distributed between GPUs and the data-parallel GPUs process them independently.”
- Denoising diffusion
A way to generate data by starting from pure noise and removing noise step by step with a trained denoiser.
What NVIDIA says (2)
“first draw a random image of pure white noise, and then chip away at the noise level”
“Training a denoiser network (typically a U-Net) with basic loss—the mean square error between its output and the clean target—achieves precisely this result.”
- DreamBooth
A fine-tuning method that teaches a text-to-image model a specific subject from a few images, bound to a unique identifier.
What NVIDIA says (2)
“you only need a few images of a specific subject to fine-tune a pretrained text-to-image model”
“so that it learns to bind a unique identifier with a special subject.”
- Dynamic batching
Combining separate inference requests into batches on the server to raise throughput.
What NVIDIA says (2)
“Dynamic batching is a feature of Triton that allows inference requests to be combined by the server, so that a batch is created dynamically. Creating a batch of requests typically results in increased throughput.”
“By default the dynamic batcher will create batches as large as possible up to the maximum batch size and will not delay when forming batches.”
E
- Embedding
A list of numbers that represents the meaning of an input such as text or an image.
What NVIDIA says (2)
“refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
“Image embeddings are the backbone to many data curation operations in NeMo Curator.”
- Estimator
A scikit-learn algorithm object that fits data to produce a model.
What NVIDIA says (1)
“An estimator is the core machine learning algorithm that fits the training data to produce a model.”
- Experiment Manager
The NeMo component that sets up logging and checkpoints for each training run.
What NVIDIA says (2)
“The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”
“resume training if checkpoints already exist resume_if_exists : True”
- Explainable AI
Tools and techniques that help people understand why a model made a decision and how it works.
What NVIDIA says (2)
“is a set of tools and techniques used by organizations to help people better understand why a model makes certain decisions and how it works.”
“Techniques with names like LIME and SHAP offer very literal mathematical answers to this question”
- Exploratory data analysis
A first look at data to learn its shape, gaps and patterns.
What NVIDIA says (1)
“However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”
F
- Feature engineering
Shaping raw data into useful model inputs, for example by scaling or encoding.
What NVIDIA says (2)
“Transformers apply algorithms to clean or reshape the training data before it is fed into a model.”
“For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”
- Federated learning
Training a shared model across sites by sharing model updates, while the data stays at each site.
What NVIDIA says (1)
“Federated learning is a way to develop and validate AI models from diverse data sources while mitigating the risk of compromising data security or privacy, as the data never leaves individual sites.”
- Few-shot prompting
Putting a few worked examples in the prompt to show the model what to do.
What NVIDIA says (1)
“Do this by providing some few-shot examples, where the reasoning process is explained. When the LLM answers the prompt, it shows its reasoning process as well.”
- FID-CLIP evaluation
Image-generation scoring that pairs FID, a realism measure, with a CLIP score for prompt match.
What NVIDIA says (1)
“Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”
- Freezing
Keeping some model weights fixed during training so only other parts learn.
What NVIDIA says (1)
“freeze : If set to True , the model parameters will not be updated during training.”
H
- Hallucination
A plausible but incorrect answer from a generative model.
What NVIDIA says (1)
“It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”
- Heat map
A grid chart colored by a value across two categories.
What NVIDIA says (1)
“An hvPlot heat map showing trips by hour and day of week, per month”
- Histogram
A chart of how many values fall into each range.
What NVIDIA says (2)
“An hvPlot histogram of trip durations generated with the Divvy dataset”
“In this instance, the vast majority of bike trips appear under 20 minutes.”
- Hydra
A configuration tool that merges YAML files and command-line overrides; NeMo uses it.
What NVIDIA says (2)
“NeMo uses Hydra for configuring both NeMo models and the PyTorch Lightning Trainer.”
“Configuration with Hydra always has the following precedence CLI > YAML > Dataclass.”
- Hyperparameter
A setting chosen before training, such as learning rate or LoRA rank.
What NVIDIA says (2)
“Optimizers and learning rate schedules are configurable across all NeMo models and have their own namespace.”
“For effective model selection, scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”
I
- Imagen
A cascaded text-to-image diffusion model that generates at 64x64 and then upsamples to 256x256 and 1024x1024.
What NVIDIA says (2)
“Imagen first generates an image at a 64x64 resolution and then upsamples the generated image to 256x256 and 1024x1024 resolutions, all using diffusion models.”
“Imagen employs a text encoder, typically T5, to encode textual features.”
- Import guarding
Writing code so an optional package is used only when it is installed.
What NVIDIA says (1)
“In either of these cases, it’s important to guard the optional imports.”
L
- Latent space
A compressed representation of data in which a model can work more cheaply than on raw pixels.
What NVIDIA says (2)
“an input image is condensed from 512x512x3 dimensions to 64x64x4.”
“This compression results in decreased memory and computational requirements when compared to pixel-space diffusion models.”
- LLM-as-a-judge
Using one large language model to grade another model's outputs.
What NVIDIA says (2)
“This method excels for tasks where automated metrics fall short, such as assessing coherence and creativity.”
“it’s important to note that LLM-as-a-judge evaluations may introduce biases inherent in the LLM evaluator training data.”
- Loss scaling
Multiplying the loss before backpropagation so small FP16 gradients do not round to zero.
What NVIDIA says (2)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
“If an overflow occurs, skip the weight update and decrease the scaling factor.”
- Low-Rank Adaptation
A PEFT method that trains small low-rank matrices; its rank r is a hyperparameter.
What NVIDIA says (2)
“is a hyperparameter that controls the rank of the decomposition”
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”
M
- Mean absolute error
The average size of prediction errors, which treats all errors equally.
What NVIDIA says (1)
“All errors are treated equally, so the metric is robust to outliers.”
- Memory-bound
An operation limited by data movement rather than math speed.
What NVIDIA says (1)
“speeding up calculation does not improve performance.”
- Mixed precision
Training mostly in 16-bit formats while keeping key values in 32-bit, to save memory and time.
What NVIDIA says (2)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
“Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32). Lowering the required memory enables training of larger models or training with larger mini-batches.”
- MLOps
Practices for building, deploying and tracking AI models reliably.
What NVIDIA says (2)
“AI models require careful tracking through cycles of experiments, tuning and retraining.”
“Align stakeholders on the goals, create an organizational structure that defines who owns what, then define responsibilities and roles”
- Modality
One type of data, such as text, image, audio or video.
What NVIDIA says (2)
“Video NeVa adds support for video modality in NeVa by representing video as multiple image frames.”
“Embed all modalities into the same vector space Ground all modalities into one primary modality Have separate stores for different modalities”
- Modality adapter
A module that maps embeddings from one modality into the embedding space of a language model.
What NVIDIA says (1)
“A modality adapter that processes the audio embeddings and produces a sequence of embeddings in the same latent space as the token embeddings of a pretrained LLM.”
- Model card
A short document that describes a model, its intended use and its limits for developers and users.
What NVIDIA says (2)
“Four subsections detailing model-specific information concerning Bias, Explainability, Privacy, and Safety and Security.”
“They are built to communicate to nontechnical stakeholders, customers using the software, and even students.”
- Multimodal model
A model that works with more than one type of data, such as text, images, audio or video.
What NVIDIA says (2)
“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”
“Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”
N
- NeMo Curator
NVIDIA's data curation library with classifiers, deduplication, quality filters and PII removal.
What NVIDIA says (2)
“Image embeddings are the backbone to many data curation operations in NeMo Curator.”
“Aesthetic classifiers can be used to assess the subjective quality of an image.”
- NeMo Guardrails
An open-source package for programmable rules that check what goes into and out of an LLM app.
What NVIDIA says (2)
“Use it to block, alter, or validate unsafe, off-topic, malicious, or policy-violating user inputs and model responses.”
“is an open-source Python package for adding programmable guardrails to LLM-based applications.”
- NeVA
NVIDIA's NeMo vision-language model that joins a large language model to a CLIP vision encoder through a projection layer.
What NVIDIA says (2)
“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”
“The encoder retrieves visual features from images and intertwines them with language embeddings using a modifiable projection matrix.”
- NSFW classifier
A NeMo Curator model that scores how likely an image is unsafe, from 0 to 1.
What NVIDIA says (1)
“a value between 0 and 1 where 1 means the content is NSFW.”
- NumPy
The core Python library for arrays and linear algebra.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
- NVIDIA Cosmos
An NVIDIA platform of world foundation models for building Physical AI systems.
What NVIDIA says (1)
“NVIDIA Cosmos is a developer-first platform for designing Physical AI systems.”
- NVIDIA DALI
A GPU-accelerated library for loading and preprocessing image, video and audio data.
What NVIDIA says (2)
“The NVIDIA Data Loading Library (DALI) is a GPU-accelerated library for data loading and pre-processing to accelerate deep learning applications.”
“It provides a collection of highly optimized building blocks for loading and processing image, video and audio data.”
- NVIDIA NIM
Containerized NVIDIA inference microservices; NIM for VLMs exposes an OpenAI-compatible API.
What NVIDIA says (2)
“NIM VLM exposes an OpenAI-compatible inference API backed by vLLM”
“GET /v1/health/ready Readiness probe. Returns 200 when the model is loaded and inference is available.”
- NVIDIA Riva
An NVIDIA SDK for GPU-accelerated speech recognition and speech synthesis.
What NVIDIA says (2)
“Automatic Speech Recognition (ASR) takes an audio stream or audio buffer as input and returns one or more text transcripts”
“based on a two-stage pipeline.”
- NVIDIA TAO
An NVIDIA toolkit for fine-tuning pretrained vision models on your data and exporting them to ONNX.
What NVIDIA says (2)
“You can select from 100+ pre-trained vision AI models on NGC and fine-tune them on your own dataset”
“TAO outputs trained models in ONNX format”
P
- pandas
The most popular Python library for working with tables of data.
What NVIDIA says (2)
“pandas is the most popular software library for data manipulation and data analysis for the Python programming languages.”
“pandas is designed to run on a single core and starts slowing down when data size hits 1-2 GB”
- Parameter-efficient fine-tuning
Adapting a large model by training small added modules while the base model stays frozen.
What NVIDIA says (1)
“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”
- Perplexity
A measure of how uncertain a language model is when predicting text; lower is better.
What NVIDIA says (1)
“Quantifies uncertainty in predicting sequences of words, with lower values indicating better predictive performance.”
- Personal identifiable information
Data that can identify a person, such as a name, email or phone number.
What NVIDIA says (1)
“The purpose of the personal identifiable information (PII) de-identification tool is to help scrub sensitive data out of datasets.”
- Precision@K
The share of the top K results that are relevant.
What NVIDIA says (1)
“Measures the proportion of retrieved documents that are relevant in a grouping of K retrieved documents.”
- Prior preservation loss
A DreamBooth loss term that uses the model's own samples to stop language drift and loss of variety.
What NVIDIA says (2)
“problems like language drift and decreased output variety often arise.”
“it guides the model using its self-generated samples and incorporates the discrepancy between the model-predicted noise on these samples.”
- Projection layer
A learned layer that maps vectors from one embedding space into another, such as image features into text embeddings.
What NVIDIA says (2)
“The encoder retrieves visual features from images and intertwines them with language embeddings using a modifiable projection matrix.”
“Transitioning from a linear to a dual-layer MLP projection markedly bolsters LLaVA-1.5’s multimodal faculties”
- Prompt engineering
Writing model inputs to get the output you want.
What NVIDIA says (2)
“Zero-shot means prompting the model without any example of expected behavior from the model.”
“Do this by providing some few-shot examples, where the reasoning process is explained. When the LLM answers the prompt, it shows its reasoning process as well.”
- PyTorch Lightning
A PyTorch library that runs the training loop; NeMo builds on it.
What NVIDIA says (1)
“NeMo uses Hydra for configuring both NeMo models and the PyTorch Lightning Trainer.”
Q
- Quantization
Running a model with fewer bits, such as INT8, to cut latency and energy.
What NVIDIA says (1)
“The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”
R
- R-squared
The share of variance in the target that a model explains.
What NVIDIA says (1)
“R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”
- RAPIDS cuDF
A GPU dataframe library with a pandas-like API.
What NVIDIA says (1)
“For that middle ground of 2-10 GB, RAPIDS cuDF is the Goldilocks solution that is just right.”
- Readiness probe
A health check that tells an orchestrator when a service can take traffic.
What NVIDIA says (1)
“GET /v1/health/ready Readiness probe. Returns 200 when the model is loaded and inference is available.”
- Recall@K
The share of all relevant items that appear in the top K results.
What NVIDIA says (1)
“Assesses the proportion of relevant documents that are successfully retrieved in a grouping of K retrieved documents.”
- Residual (skip) connection
A path that adds a block's input to its output, so gradients can flow through deep networks.
What NVIDIA says (1)
“ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”
- Retrieval-augmented generation
A method that retrieves relevant data at query time and gives it to the model as context.
What NVIDIA says (2)
“Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”
“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
- Root mean squared error
The square root of the mean squared error, on the same scale as the target.
What NVIDIA says (1)
“an RMSE of 10 does not actually mean you are off by 10 units on average.”
- ROUGE
A text-overlap score from 0 to 1 between generated and reference summaries.
What NVIDIA says (2)
“Measures overlap between machine-generated and human-generated summaries.”
“The ROUGE score ranges between 0 and 1, with higher scores indicating higher similarity.”
S
- scikit-learn
A Python library for traditional machine learning built on NumPy.
What NVIDIA says (2)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
“An estimator is the core machine learning algorithm that fits the training data to produce a model.”
- Semantic deduplication
Removing data pairs that mean nearly the same thing, found by clustering embeddings and comparing cosine similarity.
What NVIDIA says (2)
“uses embeddings to identify and remove “semantic duplicates” - data pairs that are semantically similar but not exactly identical.”
“Data pairs with cosine similarity above a threshold are considered semantic duplicates.”
- Sequence packing
Joining several short training examples into one long sequence to avoid padding.
What NVIDIA says (2)
“Many sequences are short, and a few are very long, conforming to Zipf’s Law.”
“Sequence packing is a training technique wherein multiple training sequences (examples) are concatenated into one long sequence (pack).”
- spaCy
A Python library for natural-language processing, such as named-entity recognition.
What NVIDIA says (1)
“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”
- SpeechLLM
A large language model that also takes audio, through an audio encoder and a modality adapter.
What NVIDIA says (2)
“A modality adapter that processes the audio embeddings and produces a sequence of embeddings in the same latent space as the token embeddings of a pretrained LLM.”
“One way to incorporate speech into an LLM is to concatenate speech features with the token embeddings of the input text prompt before feeding them into the LLM.”
- Stable Diffusion
A latent text-to-image diffusion model made of a U-Net, a VAE image encoder and a CLIP text encoder.
What NVIDIA says (1)
“Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”
- Stakeholder
Anyone affected by an AI project, such as the client, users and operators.
What NVIDIA says (1)
“Align stakeholders on the goals, create an organizational structure that defines who owns what, then define responsibilities and roles”
- Synthetic data
Generated data used to fill gaps in real data, including to reduce bias.
What NVIDIA says (1)
“Synthetic datasets offer one solution to reduce unwanted bias in training data”
- System prompt
Developer-written instructions sent with every user prompt to set the model's behavior.
What NVIDIA says (1)
“This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”
T
- T5
A transformer text encoder that Imagen typically uses for prompts.
What NVIDIA says (1)
“Imagen employs a text encoder, typically T5, to encode textual features.”
- Temperature
A sampling setting that controls how random the next-token choice is.
What NVIDIA says (1)
“Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”
- Tensor Core
A GPU unit for fast matrix math that works best when layer sizes are aligned.
What NVIDIA says (1)
“Tensor Cores are most efficient when key parameters of the operation are multiples of 4 if using TF32, 8 if using FP16, or 16 if using INT8”
- Text encoder
A network that turns a prompt into embeddings that guide a generative model.
What NVIDIA says (2)
“The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”
“Imagen employs a text encoder, typically T5, to encode textual features.”
- Text-to-speech
Turning text into spoken audio.
What NVIDIA says (1)
“based on a two-stage pipeline.”
- Top-k sampling
Sampling the next token only from the k most likely tokens.
What NVIDIA says (1)
“Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”
- Transfer learning
Reusing a model trained on one task as the start for another task.
What NVIDIA says (2)
“This deep learning technique enables developers to harness a neural network used for one task and apply it to another domain.”
“Transfer learning is useful when you have insufficient data for a new domain”
- Triton Inference Server
NVIDIA's open-source server for serving trained models.
What NVIDIA says (2)
“Dynamic batching is a feature of Triton that allows inference requests to be combined by the server, so that a batch is created dynamically. Creating a batch of requests typically results in increased throughput.”
“tritonserver --model-repository=/models”
- Trustworthy AI
An approach to AI development that puts safety and transparency first.
What NVIDIA says (2)
“Trustworthy AI is an approach to AI development that prioritizes safety and transparency for the people who interact with it.”
“They’re also transparent — providing information such as accuracy benchmarks or a description of the training dataset”
U
- U-Net
A convolutional network with a down-sampling path and an up-sampling path joined by skip connections, used as the denoiser in diffusion models.
What NVIDIA says (2)
“The Unet processes the noisy latents (x) to predict the noise, utilizing a conditional model which also incorporates the timestep (t) and text embedding for guidance.”
“Scaling skip connection by 1/sqrt(2)”
- Unwanted bias
Systematic unfairness in model outputs, often from narrow training data.
What NVIDIA says (2)
“AI models are trained by humans, often using data that is limited by size, scope and diversity.”
“trustworthy AI developers mitigate potential unwanted bias by looking for clues and patterns that suggest an algorithm is discriminatory”
V
- Variational autoencoder
An encoder-decoder network that compresses images into a smaller latent space and rebuilds them.
What NVIDIA says (2)
“Subsequently, during inference, the decoder reverses this process by transforming denoised latent representations back into their original, tangible image forms.”
“an input image is condensed from 512x512x3 dimensions to 64x64x4.”
- Vector database
A store of embeddings that can find the vectors closest to a query.
What NVIDIA says (2)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
“refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
- Vision encoder
A network that turns an image into feature vectors that other parts of a model can use.
What NVIDIA says (2)
“It adeptly fuses large language-centric models, such as NVGPT or LLaMA, with a vision encoder.”
“NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”
- Vision language model
A generative model that combines a large language model with a vision encoder so it can answer questions about images and video.
What NVIDIA says (2)
“Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”
“Out of the box, VLMs have strong zero-shot performance on a variety of vision tasks”
- Vision Transformer
A transformer that splits an image into patches and processes them like tokens.
What NVIDIA says (1)
“CLIP’s vision model is based on the Vision Transformer (ViT) architecture.”
W
- WebDataset
A data format that packs training samples into tar files, called shards, that stream well during training.
What NVIDIA says (1)
“The pipeline processes the dataset into the WebDataset format, consisting of tar files of equal sizes for efficient training.”
X
- XGBoost
A library for gradient-boosted decision trees.
What NVIDIA says (1)
“is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”
Z
- Zero-shot
Doing a task with no examples given in the prompt or for training.
What NVIDIA says (2)
“Zero-shot means prompting the model without any example of expected behavior from the model.”
“Out of the box, VLMs have strong zero-shot performance on a variety of vision tasks”