NCA-GENL glossary
The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.
A
- AI data flywheel
A loop where data from AI use improves the models, which produce better data.
What NVIDIA says (1)
“An AI data flywheel is a self-improving loop”
- Attention
A mechanism that lets a model weigh how much each element of the input depends on every other element.
What NVIDIA says (1)
“to detect subtle ways even distant data elements in a series influence and depend on each other”
B
- Bagging / boosting
Tree-ensemble methods: bagging (random forest) cuts variance; boosting (GBDT) cuts bias.
What NVIDIA says (1)
“Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”
- Byte Pair Encoding
A subword tokenizer that starts from characters and repeatedly merges the most frequent adjacent pairs.
What NVIDIA says (1)
“BPE starts with character vocabulary and iteratively merges frequent adjacent character pairs into new vocabulary terms”
C
- Catastrophic forgetting
When a model loses earlier knowledge while learning new data.
What NVIDIA says (1)
“the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”
- Chain-of-thought prompting
Prompting the model with worked reasoning so it shows its reasoning steps.
What NVIDIA says (1)
“Do this by providing some few-shot examples, where the reasoning process is explained.”
- Chunking
Breaking long documents into smaller pieces before embedding and retrieval.
What NVIDIA says (1)
“A chunking strategy is the method of breaking down large documents into smaller, manageable pieces for AI retrieval.”
- Confidential computing
Protecting data and models while in use inside hardware Trusted Execution Environments.
What NVIDIA says (1)
“hardware-rooted Trusted Execution Environments (TEEs)”
- Context window
The maximum amount of text (prompt plus retrieved context) an LLM can take in at once.
What NVIDIA says (1)
“The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”
- Cosine similarity
A similarity score based on the angle between two vectors.
What NVIDIA says (1)
“Cosine similarity: Focuses on the angle between vectors.”
- Cross-validation
Splitting training data into subsets to train and validate in turns.
What NVIDIA says (1)
“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”
- cuDF
RAPIDS GPU DataFrame library with a pandas-like API.
What NVIDIA says (1)
“enables GPU acceleration that unlocks access to your data insights through a familiar pandas-like API.”
- cuML
RAPIDS GPU library for traditional machine learning, including text vectorizers.
What NVIDIA says (1)
“subpackage in cuML by adding Count and TF-IDF vectorizer”
- cuxfilter
RAPIDS library for GPU-accelerated cross-filtering dashboards.
What NVIDIA says (1)
“cuxfilter enables GPU accelerated cross-filtering dashboards from notebooks”
D
- Decontamination
Removing test data that leaked into training data.
What NVIDIA says (1)
“the potential leakage of test data into training datasets”
- Deduplication (exact / fuzzy / semantic)
Removing duplicate or near-duplicate documents from training data.
What NVIDIA says (1)
“Fuzzy deduplication addresses near-duplicate content using MinHash signatures and Locality-Sensitive Hashing (LSH)”
E
- Embedding
A list of numbers (a vector) that represents the meaning of text, so similar text gets similar vectors.
What NVIDIA says (1)
“Generating embeddings involves converting data into high-dimensional vectors, which represent text in a numerical format.”
- Explainable AI
Tools and techniques that help people understand why a model made a decision.
What NVIDIA says (1)
“help people better understand why a model makes certain decisions and how it works.”
- Exploratory data analysis
Open-ended exploration of a dataset to understand it before modeling.
What NVIDIA says (1)
“exploratory data analysis (EDA)”
F
- Federated learning
Training across many sites while the data stays at each site.
What NVIDIA says (1)
“as the data never leaves individual sites.”
- Few-shot prompt
A prompt that includes a few example prompt-and-answer pairs.
What NVIDIA says (1)
“This approach requires prepending a few sample prompts and completion pairs to the prompt”
- Foundation model
A large model trained on massive unlabeled data that can be adapted to many tasks.
What NVIDIA says (1)
“Foundation models are AI neural networks trained on massive unlabeled datasets to handle a wide variety of jobs”
- FP16 / FP32
16-bit and 32-bit floating-point number formats.
What NVIDIA says (1)
“Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32).”
G
- Grid search
Trying many hyperparameter combinations to find the best one.
What NVIDIA says (1)
“Grid search further automates hyperparameter optimization, testing various configurations for improved accuracy.”
H
- Hallucination
A plausible-sounding but incorrect answer from a model.
What NVIDIA says (1)
“a very plausible but incorrect answer, a phenomenon called hallucination.”
I
- In-flight batching
A serving technique that avoids making a whole batch wait for its longest request.
What NVIDIA says (1)
“There are methods to mitigate this, such as in-flight batching.”
- Intertoken latency
The average time between consecutive output tokens.
What NVIDIA says (1)
“is the average time between the generation of consecutive tokens in a sequence.”
K
- KV cache
Stored attention keys and values for earlier tokens; with model weights it dominates inference memory.
What NVIDIA says (1)
“the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”
L
- Large language model
A deep-learning model trained on huge amounts of text that can understand and generate language.
What NVIDIA says (1)
“Large language models (LLMs) are deep learning algorithms that can recognize, summarize, translate, predict, and generate content using very large datasets.”
- LLM agent
A system that uses an LLM to reason, plan and act with tools.
What NVIDIA says (1)
“a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”
- LLM-as-a-judge
Using one LLM to grade another LLM's answers.
What NVIDIA says (1)
“Using LLMs to assess other LLMs can introduce biases that skew results”
- Loss scaling
Scaling the loss in mixed-precision training so small gradients are not lost.
What NVIDIA says (1)
“mixed precision with loss scaling (green) matches the single precision model (black).”
- Low-Rank Adaptation
A PEFT method that trains small low-rank matrices in each layer while the original weights stay frozen.
What NVIDIA says (1)
“only trains these matrices while keeping the original LLM weights frozen.”
M
- Mixed-precision training
Training that does most math in 16-bit floats while keeping key values in 32-bit.
What NVIDIA says (1)
“by performing operations in half-precision format, while storing minimal information in single-precision”
- Mixture of experts
A model made of many specialized sub-models (experts), where a router sends each token to only a few of them.
What NVIDIA says (1)
“Mixture of experts (MoE) is an AI model architecture that uses multiple, specialized submodels, or "experts,"”
- MLOps
Best practices for running AI in production, modeled on DevOps.
What NVIDIA says (1)
“MLOps is a set of best practices for businesses to run AI successfully.”
- Model card / Model Card++
A document describing a model's capabilities; NVIDIA's version adds Bias, Explainability, Privacy, and Safety and Security sections.
What NVIDIA says (1)
“Four subsections detailing model-specific information concerning Bias, Explainability, Privacy, and Safety and Security.”
- MSE / RMSE / MAE
Regression error metrics: mean squared error, its square root, and mean absolute error.
What NVIDIA says (2)
“is simply the square root of the latter.”
“Similar to MSE and RMSE, MAE is also scale-dependent”
N
- Named-entity recognition
Finding and labeling entities in text, such as people, places and companies.
What NVIDIA says (1)
“is the task of detecting and classifying key information (entities) in text.”
- NCCL
A library for fast communication between GPUs, such as AllReduce.
What NVIDIA says (1)
“is a library providing inter-GPU communication primitives that are topology-aware”
- NeMo Curator
NVIDIA's GPU-accelerated toolkit for curating LLM training data.
What NVIDIA says (1)
“NeMo Curator supports data curation for model pretraining”
- NeMo Guardrails
An open-source library that adds programmable safety and topic rules around an LLM app.
What NVIDIA says (1)
“is an open-source Python package for adding programmable guardrails to LLM-based applications.”
- NeMo Retriever
NVIDIA microservices for indexing and querying data, with embedding and reranking.
What NVIDIA says (1)
“is a collection of microservices that present a single API for indexing and querying of user data.”
- NumPy
Python library for fast array and linear-algebra operations.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
- NVIDIA NeMo Framework
NVIDIA's framework to create, customize and deploy generative AI models.
What NVIDIA says (1)
“It enables users to efficiently create, customize, and deploy new generative AI models”
- NVIDIA NIM
Containerized model-serving microservices with industry-standard APIs.
What NVIDIA says (1)
“NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”
P
- pandas / DataFrame
Python library for tabular data; a DataFrame is a table of rows and columns.
What NVIDIA says (1)
“A pandas DataFrame is a two-dimensional, array-like table”
- Parameter-efficient fine-tuning
Fine-tuning that adds or updates only a few parameters or layers instead of the whole model.
What NVIDIA says (1)
“Parameter-efficient fine-tuning (PEFT) techniques use clever optimizations to selectively add and update few parameters or layers to the original LLM architecture.”
- Personally identifiable information
Data that can identify a person, such as a name or social security number.
What NVIDIA says (1)
“from direct identifiers like names and social security numbers”
- Principal component analysis
A dimensionality-reduction technique that keeps meaningful patterns with fewer variables.
What NVIDIA says (1)
“dimensionality reduction techniques, such as PCA”
- Prompt engineering
Shaping the input prompt to get the output you want, without changing the model's weights.
What NVIDIA says (1)
“Manipulates the prompt sent to the LLM but doesn’t alter the parameters of the LLM in any way.”
- Prompt tuning / p-tuning
Learning small virtual-token embeddings for a task while the LLM's own weights stay frozen.
What NVIDIA says (1)
“prompt tuning and p-tuning use virtual prompt embeddings that you can optimize by gradient descent.”
- Proximal policy optimization
The reinforcement-learning algorithm used in RLHF stage 3 to tune the model against the reward model.
What NVIDIA says (1)
“using reinforcement learning with a proximal policy optimization (PPO) algorithm.”
- PTQ / QAT
PTQ quantizes a trained model; QAT includes quantization during training for better accuracy.
What NVIDIA says (2)
“post-training quantization (PTQ)”
“QAT almost always produces better accuracy”
Q
- Quantization
Storing weights and activations in a lower-precision format, typically 8-bit integers.
What NVIDIA says (1)
“are converted from a floating-point representation to a lower-precision representation, typically using 8-bit integers.”
R
- Recall / NDCG
Retrieval metrics: recall counts relevant items found; NDCG also rewards ranking them high.
What NVIDIA says (1)
“Now let’s delve deeper into Recall and NDCG”
- Red teaming
Assessing an AI system from an attacker's point of view to find and reduce risks.
What NVIDIA says (1)
“mitigate any risks from the perspective of information security.”
- Reinforcement learning from human feedback
Aligning an LLM with human preferences using a reward model trained on human rankings.
What NVIDIA says (1)
“Reinforcement learning with human feedback (RLHF) is a customization technique that enables LLMs to achieve better alignment with human values and preferences.”
- Reranking
A second retrieval stage that re-scores candidate passages for relevance to the query.
What NVIDIA says (1)
“Re-ranking is typically used as a second stage after an initial fast retrieval step”
- Residual
Actual value minus predicted value.
What NVIDIA says (1)
“a residual is a difference between the actual value and the predicted value.”
- Retrieval-augmented generation
A pattern where relevant documents are retrieved at query time and given to the LLM as context for its answer.
What NVIDIA says (1)
“when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
- Reward model
A model trained on human-ranked responses that predicts which answer people would prefer.
What NVIDIA says (1)
“is used to train the RM to predict human preference.”
- R²
Proportion of the target's variance a model explains.
What NVIDIA says (1)
“represents the proportion of variance explained by a model.”
S
- scikit-learn
Python library for traditional machine learning.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy”
- spaCy
Python NLP library; used here for named-entity recognition on the CPU.
What NVIDIA says (1)
“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”
- Supervised fine-tuning
Fine-tuning the model's weights on labeled examples, such as instructions with desired answers.
What NVIDIA says (1)
“SFT with instructions leverages the intuition that NLP tasks can be described through natural language instructions”
- Supervised learning
Learning from labeled examples, as in classification and regression.
What NVIDIA says (1)
“Regression estimates the relationship between a target outcome label and one or more feature variables”
- System prompt
An instruction added before the user's prompt that tells the LLM how to behave.
What NVIDIA says (1)
“adding a system-level prompt in addition to the user prompt”
T
- Temperature
A sampling setting: lower values make output more predictable, higher values more varied.
What NVIDIA says (1)
“At a lower temperature, the model is more conservative and is limited to choosing tokens with higher probabilities.”
- Tensor / pipeline parallelism
Splitting one model across GPUs, by layer pieces (tensor) or by groups of layers (pipeline).
What NVIDIA says (1)
“Tensor parallelism involves sharding (horizontally) individual layers of the model”
- TF-IDF
A way to turn text into numeric feature vectors.
What NVIDIA says (1)
“by first vectorizing them using TF-IDF”
- Time to first token
How long until the first output token arrives.
What NVIDIA says (1)
“TTFT generally includes both request queuing time, prefill time, and network latency.”
- Token
A small unit of text, such as a word or part of a word, that a model processes.
What NVIDIA says (1)
“fragmenting text into smaller units known as tokens.”
- Tokenization
Splitting text into tokens and mapping them to numeric IDs the model can use.
What NVIDIA says (1)
“These extracted tokens are used to build a vocabulary index mapping tokens to numeric IDs”
- Transfer learning
Starting from a model trained on one task and adapting it to a related task.
What NVIDIA says (1)
“transfer learning harnesses a settled domain to pioneer new terrain.”
- Transformer
A neural-network architecture that uses attention to learn how parts of a sequence relate to each other.
What NVIDIA says (1)
“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention”
- Triton Inference Server
NVIDIA's inference server, which loads models from a model repository.
What NVIDIA says (1)
“Launching and maintaining Triton Inference Server revolves around the use of building model repositories.”
- Trustworthy AI
AI development that puts safety and transparency first for the people it affects.
What NVIDIA says (1)
“prioritizes safety and transparency for the people who interact with it.”
U
- Unsupervised learning
Finding patterns in data without labels, as in clustering.
What NVIDIA says (1)
“doesn’t have labeled data provided in advance”
- Unwanted bias
Systematic unfairness in a model, often from data limited in size, scope or diversity.
What NVIDIA says (1)
“often using data that is limited by size, scope and diversity.”
V
- Vector database
A database that stores embeddings and finds the ones most similar to a query.
What NVIDIA says (1)
“A vector database is an organized collection of vector embeddings”
Z
- Zero-shot prompt
A prompt with no examples of the expected behavior.
What NVIDIA says (1)
“Zero-shot means prompting the model without any example of expected behavior from the model.”