NCA-GENL glossary

The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.

A

AI data flywheel

A loop where data from AI use improves the models, which produce better data.

Objectives: 2.5

What NVIDIA says (1)

“An AI data flywheel is a self-improving loop”

— Data flywheel: What it is and how it works

Attention (self-attention)

A mechanism that lets a model weigh how much each element of the input depends on every other element.

Objectives: 1.7

What NVIDIA says (1)

“to detect subtle ways even distant data elements in a series influence and depend on each other”

— What Is a Transformer Model?

B

Bagging / boosting

Tree-ensemble methods: bagging (random forest) cuts variance; boosting (GBDT) cuts bias.

Objectives: 1.5

What NVIDIA says (1)

“Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”

— What Is XGBoost and Why Does It Matter?

Byte Pair Encoding (BPE)

A subword tokenizer that starts from characters and repeatedly merges the most frequent adjacent pairs.

Objectives: 3.2

What NVIDIA says (1)

“BPE starts with character vocabulary and iteratively merges frequent adjacent character pairs into new vocabulary terms”

— Mastering LLM Techniques: Training

C

Catastrophic forgetting

When a model loses earlier knowledge while learning new data.

Objectives: 3.4

What NVIDIA says (1)

“the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

Chain-of-thought prompting (CoT)

Prompting the model with worked reasoning so it shows its reasoning steps.

Objectives: 1.9

What NVIDIA says (1)

“Do this by providing some few-shot examples, where the reasoning process is explained.”

— An Introduction to Large Language Models: Prompt Engineering and P-Tuning

Chunking (text splitting)

Breaking long documents into smaller pieces before embedding and retrieval.

Objectives: 1.4

What NVIDIA says (1)

“A chunking strategy is the method of breaking down large documents into smaller, manageable pieces for AI retrieval.”

— Finding the Best Chunking Strategy for Accurate AI Responses

Confidential computing (TEE)

Protecting data and models while in use inside hardware Trusted Execution Environments.

Objectives: 5.3

What NVIDIA says (1)

“hardware-rooted Trusted Execution Environments (TEEs)”

— What Is Confidential Computing?

Context window

The maximum amount of text (prompt plus retrieved context) an LLM can take in at once.

Objectives: 2.2

What NVIDIA says (1)

“The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”

— Enhancing RAG Pipelines with Re-Ranking

Cosine similarity

A similarity score based on the angle between two vectors.

Objectives: 1.8

What NVIDIA says (1)

“Cosine similarity: Focuses on the angle between vectors.”

— What is a Vector Database and How Does it Work?

Cross-validation

Splitting training data into subsets to train and validate in turns.

Objectives: 1.5

What NVIDIA says (1)

“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”

— What is scikit-learn?

cuDF

RAPIDS GPU DataFrame library with a pandas-like API.

Objectives: 4.1, 2.3

What NVIDIA says (1)

“enables GPU acceleration that unlocks access to your data insights through a familiar pandas-like API.”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

cuML

RAPIDS GPU library for traditional machine learning, including text vectorizers.

Objectives: 2.6

What NVIDIA says (1)

“subpackage in cuML by adding Count and TF-IDF vectorizer”

— NLP and Text Processing with RAPIDS: Now Simpler and Faster

cuxfilter

RAPIDS library for GPU-accelerated cross-filtering dashboards.

Objectives: 4.4

What NVIDIA says (1)

“cuxfilter enables GPU accelerated cross-filtering dashboards from notebooks”

— Welcome to cuxfilter’s documentation — cuxfilter 26.06.00 documentation

D

Decontamination

Removing test data that leaked into training data.

Objectives: 3.2

What NVIDIA says (1)

“the potential leakage of test data into training datasets”

— Mastering LLM Techniques: Text Data Processing

Deduplication (exact / fuzzy / semantic) (MinHash, LSH)

Removing duplicate or near-duplicate documents from training data.

Objectives: 3.2

What NVIDIA says (1)

“Fuzzy deduplication addresses near-duplicate content using MinHash signatures and Locality-Sensitive Hashing (LSH)”

— Mastering LLM Techniques: Text Data Processing

E

Embedding (vector embedding)

A list of numbers (a vector) that represents the meaning of text, so similar text gets similar vectors.

Objectives: 1.8

What NVIDIA says (1)

“Generating embeddings involves converting data into high-dimensional vectors, which represent text in a numerical format.”

— RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

Explainable AI (XAI)

Tools and techniques that help people understand why a model made a decision.

Objectives: 5.3

What NVIDIA says (1)

“help people better understand why a model makes certain decisions and how it works.”

— What Is Explainable AI (XAI)?

Exploratory data analysis (EDA)

Open-ended exploration of a dataset to understand it before modeling.

Objectives: 4.1, 4.3

What NVIDIA says (1)

“exploratory data analysis (EDA)”

— Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

F

Federated learning

Training across many sites while the data stays at each site.

Objectives: 5.2

What NVIDIA says (1)

“as the data never leaves individual sites.”

— What Is Federated Learning?

Few-shot prompt

A prompt that includes a few example prompt-and-answer pairs.

Objectives: 1.9

What NVIDIA says (1)

“This approach requires prepending a few sample prompts and completion pairs to the prompt”

— Mastering LLM Techniques: Customization

Foundation model

A large model trained on massive unlabeled data that can be adapted to many tasks.

Objectives: 1.7

What NVIDIA says (1)

“Foundation models are AI neural networks trained on massive unlabeled datasets to handle a wide variety of jobs”

— What Are Foundation Models?

FP16 / FP32 (half / single precision)

16-bit and 32-bit floating-point number formats.

Objectives: 2.4, 3.1

What NVIDIA says (1)

“Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32).”

— Train With Mixed Precision

G

Trying many hyperparameter combinations to find the best one.

Objectives: 2.6

What NVIDIA says (1)

“Grid search further automates hyperparameter optimization, testing various configurations for improved accuracy.”

— What is scikit-learn?

H

Hallucination

A plausible-sounding but incorrect answer from a model.

Objectives: 1.3

What NVIDIA says (1)

“a very plausible but incorrect answer, a phenomenon called hallucination.”

— What Is Retrieval-Augmented Generation aka RAG

I

In-flight batching

A serving technique that avoids making a whole batch wait for its longest request.

Objectives: 1.1

What NVIDIA says (1)

“There are methods to mitigate this, such as in-flight batching.”

— Mastering LLM Techniques: Inference Optimization

Intertoken latency (ITL, TPOT)

The average time between consecutive output tokens.

Objectives: 3.3

What NVIDIA says (1)

“is the average time between the generation of consecutive tokens in a sequence.”

— LLM Inference Benchmarking: Fundamental Concepts

K

KV cache (key-value cache)

Stored attention keys and values for earlier tokens; with model weights it dominates inference memory.

Objectives: 1.1

What NVIDIA says (1)

“the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”

— Mastering LLM Techniques: Inference Optimization

L

Large language model (LLM)

A deep-learning model trained on huge amounts of text that can understand and generate language.

Objectives: 1.3, 2.2

What NVIDIA says (1)

“Large language models (LLMs) are deep learning algorithms that can recognize, summarize, translate, predict, and generate content using very large datasets.”

— What are Large Language Models?

LLM agent

A system that uses an LLM to reason, plan and act with tools.

Objectives: 2.2

What NVIDIA says (1)

“a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”

— Introduction to LLM Agents

LLM-as-a-judge

Using one LLM to grade another LLM's answers.

Objectives: 3.6

What NVIDIA says (1)

“Using LLMs to assess other LLMs can introduce biases that skew results”

— Mastering LLM Techniques: Evaluation

Loss scaling

Scaling the loss in mixed-precision training so small gradients are not lost.

Objectives: 3.1

What NVIDIA says (1)

“mixed precision with loss scaling (green) matches the single precision model (black).”

— Train With Mixed Precision

Low-Rank Adaptation (LoRA)

A PEFT method that trains small low-rank matrices in each layer while the original weights stay frozen.

Objectives: 3.1, 3.4

What NVIDIA says (1)

“only trains these matrices while keeping the original LLM weights frozen.”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

M

Mixed-precision training

Training that does most math in 16-bit floats while keeping key values in 32-bit.

Objectives: 3.1

What NVIDIA says (1)

“by performing operations in half-precision format, while storing minimal information in single-precision”

— Train With Mixed Precision

Mixture of experts (MoE)

A model made of many specialized sub-models (experts), where a router sends each token to only a few of them.

Objectives: 1.7

What NVIDIA says (1)

“Mixture of experts (MoE) is an AI model architecture that uses multiple, specialized submodels, or "experts,"”

— What Is Mixture of Experts (MoE) and How It Works?

MLOps (machine learning operations)

Best practices for running AI in production, modeled on DevOps.

Objectives: 2.5

What NVIDIA says (1)

“MLOps is a set of best practices for businesses to run AI successfully.”

— What is MLOps?

Model card / Model Card++

A document describing a model's capabilities; NVIDIA's version adds Bias, Explainability, Privacy, and Safety and Security sections.

Objectives: 5.3

What NVIDIA says (1)

“Four subsections detailing model-specific information concerning Bias, Explainability, Privacy, and Safety and Security.”

— Enhancing AI Transparency and Ethical Considerations with Model Card++

MSE / RMSE / MAE

Regression error metrics: mean squared error, its square root, and mean absolute error.

Objectives: 4.2

What NVIDIA says (2)

“is simply the square root of the latter.”

— A Comprehensive Overview of Regression Evaluation Metrics

“Similar to MSE and RMSE, MAE is also scale-dependent”

— A Comprehensive Overview of Regression Evaluation Metrics

N

Named-entity recognition (NER)

Finding and labeling entities in text, such as people, places and companies.

Objectives: 1.6

What NVIDIA says (1)

“is the task of detecting and classifying key information (entities) in text.”

— Token Classification Model with Named Entity Recognition (NER) — NVIDIA NeMo Framework User Guide

NCCL (NVIDIA Collective Communications Library)

A library for fast communication between GPUs, such as AllReduce.

Objectives: 2.4

What NVIDIA says (1)

“is a library providing inter-GPU communication primitives that are topology-aware”

— Overview of NCCL — NCCL 2.32.3 documentation

NeMo Curator

NVIDIA's GPU-accelerated toolkit for curating LLM training data.

Objectives: 3.2

What NVIDIA says (1)

“NeMo Curator supports data curation for model pretraining”

— Scale and Curate High-Quality Datasets for LLM Training with NVIDIA NeMo Curator

NeMo Guardrails (guardrails)

An open-source library that adds programmable safety and topic rules around an LLM app.

Objectives: 2.7, 5.3

What NVIDIA says (1)

“is an open-source Python package for adding programmable guardrails to LLM-based applications.”

— Overview | NVIDIA NeMo Guardrails Library Developer Guide

NeMo Retriever

NVIDIA microservices for indexing and querying data, with embedding and reranking.

Objectives: 2.4

What NVIDIA says (1)

“is a collection of microservices that present a single API for indexing and querying of user data.”

— NVIDIA NeMo Retriever

NumPy

Python library for fast array and linear-algebra operations.

Objectives: 1.6

What NVIDIA says (1)

“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”

— What is scikit-learn?

NVIDIA NeMo Framework (NeMo)

NVIDIA's framework to create, customize and deploy generative AI models.

Objectives: 2.4

What NVIDIA says (1)

“It enables users to efficiently create, customize, and deploy new generative AI models”

— Overview — NVIDIA NeMo Framework User Guide

NVIDIA NIM (NVIDIA Inference Microservices)

Containerized model-serving microservices with industry-standard APIs.

Objectives: 2.4

What NVIDIA says (1)

“NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”

— Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

P

pandas / DataFrame

Python library for tabular data; a DataFrame is a table of rows and columns.

Objectives: 2.3, 4.3

What NVIDIA says (1)

“A pandas DataFrame is a two-dimensional, array-like table”

— What Is Pandas and Why Does it Matter?

Parameter-efficient fine-tuning (PEFT)

Fine-tuning that adds or updates only a few parameters or layers instead of the whole model.

Objectives: 3.1, 3.4

What NVIDIA says (1)

“Parameter-efficient fine-tuning (PEFT) techniques use clever optimizations to selectively add and update few parameters or layers to the original LLM architecture.”

— Mastering LLM Techniques: Customization

Personally identifiable information (PII)

Data that can identify a person, such as a name or social security number.

Objectives: 3.2, 5.2

What NVIDIA says (1)

“from direct identifiers like names and social security numbers”

— Mastering LLM Techniques: Text Data Processing

Principal component analysis (PCA)

A dimensionality-reduction technique that keeps meaningful patterns with fewer variables.

Objectives: 1.10

What NVIDIA says (1)

“dimensionality reduction techniques, such as PCA”

— What is scikit-learn?

Prompt engineering

Shaping the input prompt to get the output you want, without changing the model's weights.

Objectives: 1.9

What NVIDIA says (1)

“Manipulates the prompt sent to the LLM but doesn’t alter the parameters of the LLM in any way.”

— Mastering LLM Techniques: Customization

Prompt tuning / p-tuning (prompt learning)

Learning small virtual-token embeddings for a task while the LLM's own weights stay frozen.

Objectives: 3.4

What NVIDIA says (1)

“prompt tuning and p-tuning use virtual prompt embeddings that you can optimize by gradient descent.”

— Mastering LLM Techniques: Customization

Proximal policy optimization (PPO)

The reinforcement-learning algorithm used in RLHF stage 3 to tune the model against the reward model.

Objectives: 3.5

What NVIDIA says (1)

“using reinforcement learning with a proximal policy optimization (PPO) algorithm.”

— Mastering LLM Techniques: Customization

PTQ / QAT (post-training quantization / quantization-aware training)

PTQ quantizes a trained model; QAT includes quantization during training for better accuracy.

Objectives: 3.4

What NVIDIA says (2)

“post-training quantization (PTQ)”

— Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

“QAT almost always produces better accuracy”

— Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

Q

Quantization

Storing weights and activations in a lower-precision format, typically 8-bit integers.

Objectives: 3.4

What NVIDIA says (1)

“are converted from a floating-point representation to a lower-precision representation, typically using 8-bit integers.”

— Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

R

Recall / NDCG (normalized discounted cumulative gain)

Retrieval metrics: recall counts relevant items found; NDCG also rewards ranking them high.

Objectives: 3.6

What NVIDIA says (1)

“Now let’s delve deeper into Recall and NDCG”

— Evaluating Retriever for Enterprise-Grade RAG

Red teaming

Assessing an AI system from an attacker's point of view to find and reduce risks.

Objectives: 3.3

What NVIDIA says (1)

“mitigate any risks from the perspective of information security.”

— NVIDIA AI Red Team: An Introduction

Reinforcement learning from human feedback (RLHF)

Aligning an LLM with human preferences using a reward model trained on human rankings.

Objectives: 3.5

What NVIDIA says (1)

“Reinforcement learning with human feedback (RLHF) is a customization technique that enables LLMs to achieve better alignment with human values and preferences.”

— Mastering LLM Techniques: Customization

Reranking (re-ranking)

A second retrieval stage that re-scores candidate passages for relevance to the query.

Objectives: 2.2

What NVIDIA says (1)

“Re-ranking is typically used as a second stage after an initial fast retrieval step”

— Enhancing RAG Pipelines with Re-Ranking

Residual

Actual value minus predicted value.

Objectives: 4.2

What NVIDIA says (1)

“a residual is a difference between the actual value and the predicted value.”

— A Comprehensive Overview of Regression Evaluation Metrics

Retrieval-augmented generation (RAG)

A pattern where relevant documents are retrieved at query time and given to the LLM as context for its answer.

Objectives: 1.3, 1.4, 2.2

What NVIDIA says (1)

“when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”

— RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

Reward model (RM)

A model trained on human-ranked responses that predicts which answer people would prefer.

Objectives: 3.5

What NVIDIA says (1)

“is used to train the RM to predict human preference.”

— Mastering LLM Techniques: Customization

R² (coefficient of determination)

Proportion of the target's variance a model explains.

Objectives: 4.2

What NVIDIA says (1)

“represents the proportion of variance explained by a model.”

— A Comprehensive Overview of Regression Evaluation Metrics

S

scikit-learn (sklearn)

Python library for traditional machine learning.

Objectives: 1.10, 2.6

What NVIDIA says (1)

“scikit-learn is a versatile Python library built on NumPy”

— What is scikit-learn?

spaCy

Python NLP library; used here for named-entity recognition on the CPU.

Objectives: 1.6, 2.3

What NVIDIA says (1)

“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”

— Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

Supervised fine-tuning (SFT)

Fine-tuning the model's weights on labeled examples, such as instructions with desired answers.

Objectives: 3.1

What NVIDIA says (1)

“SFT with instructions leverages the intuition that NLP tasks can be described through natural language instructions”

— Mastering LLM Techniques: Customization

Supervised learning

Learning from labeled examples, as in classification and regression.

Objectives: 1.5

What NVIDIA says (1)

“Regression estimates the relationship between a target outcome label and one or more feature variables”

— What is Machine Learning and Why Does It Matter?

System prompt

An instruction added before the user's prompt that tells the LLM how to behave.

Objectives: 1.9

What NVIDIA says (1)

“adding a system-level prompt in addition to the user prompt”

— Mastering LLM Techniques: Customization

T

Temperature

A sampling setting: lower values make output more predictable, higher values more varied.

Objectives: 1.9

What NVIDIA says (1)

“At a lower temperature, the model is more conservative and is limited to choosing tokens with higher probabilities.”

— How to Get Better Outputs from Your Large Language Model

Tensor / pipeline parallelism (model parallelism)

Splitting one model across GPUs, by layer pieces (tensor) or by groups of layers (pipeline).

Objectives: 1.1

What NVIDIA says (1)

“Tensor parallelism involves sharding (horizontally) individual layers of the model”

— Mastering LLM Techniques: Inference Optimization

TF-IDF (term frequency–inverse document frequency)

A way to turn text into numeric feature vectors.

Objectives: 2.6

What NVIDIA says (1)

“by first vectorizing them using TF-IDF”

— NLP and Text Processing with RAPIDS: Now Simpler and Faster

Time to first token (TTFT)

How long until the first output token arrives.

Objectives: 2.1, 3.3

What NVIDIA says (1)

“TTFT generally includes both request queuing time, prefill time, and network latency.”

— LLM Inference Benchmarking: Fundamental Concepts

Token

A small unit of text, such as a word or part of a word, that a model processes.

Objectives: 3.2

What NVIDIA says (1)

“fragmenting text into smaller units known as tokens.”

— Mastering LLM Techniques: Training

Tokenization

Splitting text into tokens and mapping them to numeric IDs the model can use.

Objectives: 3.2

What NVIDIA says (1)

“These extracted tokens are used to build a vocabulary index mapping tokens to numeric IDs”

— Mastering LLM Techniques: Training

Transfer learning

Starting from a model trained on one task and adapting it to a related task.

Objectives: 3.1

What NVIDIA says (1)

“transfer learning harnesses a settled domain to pioneer new terrain.”

— What Is Transfer Learning?

Transformer

A neural-network architecture that uses attention to learn how parts of a sequence relate to each other.

Objectives: 1.7

What NVIDIA says (1)

“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention”

— What Is a Transformer Model?

Triton Inference Server (Triton)

NVIDIA's inference server, which loads models from a model repository.

Objectives: 2.7

What NVIDIA says (1)

“Launching and maintaining Triton Inference Server revolves around the use of building model repositories.”

— Quickstart — NVIDIA Triton Inference Server

Trustworthy AI

AI development that puts safety and transparency first for the people it affects.

Objectives: 5.1

What NVIDIA says (1)

“prioritizes safety and transparency for the people who interact with it.”

— What Is Trustworthy AI?

U

Unsupervised learning

Finding patterns in data without labels, as in clustering.

Objectives: 1.2, 1.5

What NVIDIA says (1)

“doesn’t have labeled data provided in advance”

— What is Machine Learning and Why Does It Matter?

Unwanted bias

Systematic unfairness in a model, often from data limited in size, scope or diversity.

Objectives: 5.4

What NVIDIA says (1)

“often using data that is limited by size, scope and diversity.”

— What Is Trustworthy AI?

V

Vector database (vector DB)

A database that stores embeddings and finds the ones most similar to a query.

Objectives: 1.6, 1.4

What NVIDIA says (1)

“A vector database is an organized collection of vector embeddings”

— What is a Vector Database and How Does it Work?

Z

Zero-shot prompt

A prompt with no examples of the expected behavior.

Objectives: 1.9

What NVIDIA says (1)

“Zero-shot means prompting the model without any example of expected behavior from the model.”

— An Introduction to Large Language Models: Prompt Engineering and P-Tuning