Experimentation
22% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI
3.1 Training and training optimization
Techniques that make training faster or cheaper.
Key points
Mixed precision runs operations in FP16 (half precision) while keeping minimal information in FP32 (single precision). NVIDIA's training curves show mixed precision without loss scaling diverging, while with loss scaling it matches the single-precision model.
What NVIDIA says (2)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
“by performing operations in half-precision format, while storing minimal information in single-precision”
Low-rank adaptation (LoRA) belongs to the parameter-efficient fine-tuning (PEFT) family. It inserts small low-rank matrices into each layer and trains only those, keeping the original large language model (LLM) weights frozen.
What NVIDIA says (2)
“LoRA is a fine-tuning method that introduces low-rank matrices into each layer of the LLM architecture, and only trains these matrices while keeping the original LLM weights frozen.”
“LoRA tuning is a type of tuning family called Parameter Efficient Fine-Tuning (PEFT).”
Transfer learning starts from an existing trained model. NVIDIA's example deletes the final “loss output” layer, adds a new one for the new task, then trains on the smaller new dataset: the whole network, the last few layers, or just that layer.
What NVIDIA says (2)
“Next, you would take your smaller dataset for horses and train it on the entire 50-layer neural network or the last few layers or just the loss layer alone.”
“First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”
The rank r is a low-rank adaptation (LoRA) hyperparameter that controls the rank of the decomposition. NVIDIA says a smaller r saves parameters and memory and trains faster, but can capture less task-specific information.
What NVIDIA says (2)
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”
“is a hyperparameter that controls the rank of the decomposition”
Key terms: Parameter-efficient fine-tuning Low-Rank Adaptation Supervised fine-tuning Mixed-precision training FP16 / FP32 Loss scaling Transfer learning
Try it: LoRA lab
3.2 Preparing large datasets
Cleaning, deduplicating and tokenizing data for pretraining and fine-tuning.
Key points
Tokenization splits text into smaller units called tokens. NVIDIA explains that word tokenizers create huge vocabularies and out-of-vocabulary tokens, while character tokenizers create long sequences. Subword tokenizers split rare words into meaningful subwords. Popular ones are byte-pair encoding (BPE), WordPiece, Unigram and SentencePiece.
What NVIDIA says (5)
“Word-based tokenizers lead to a large vocabulary size and words not seen during the tokenizer training process cause many out-of-vocabulary tokens.”
“Popular subword tokenization algorithms include Byte Pair Encoding (BPE), WordPiece, Unigram, and SentencePiece.”
“Character-based tokenizers lead to long sequences”
“The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”
“the process of tokenization assumes a pivotal role in fragmenting text into smaller units known as tokens.”
Exact deduplication hashes each document and keeps one per hash bucket, so it only removes identical copies. Fuzzy deduplication finds near-duplicates using MinHash signatures and Locality-Sensitive Hashing (LSH).
What NVIDIA says (2)
“Fuzzy deduplication addresses near-duplicate content using MinHash signatures and Locality-Sensitive Hashing (LSH) to identify similar documents.”
“This method generates hash signatures for each document and groups documents by their hashes into buckets”
NVIDIA's data-processing guide says deduplication improves training efficiency, cuts compute cost and keeps data diverse. It also helps prevent overfitting to repeated content.
What NVIDIA says (2)
“Deduplication is essential for improving model training efficiency, reducing computational costs, and ensuring data diversity.”
“It helps prevent models from overfitting to repeated content and improves generalization.”
NVIDIA lists Unicode fixing and language identification as crucial early steps when curating large web-scraped text corpora.
What NVIDIA says (1)
“Unicode fixing and language identification represent crucial early steps in the data curation pipeline, particularly when dealing with large-scale web-scraped text corpora.”
Heuristic filtering uses rule-based metrics and statistics to remove low-quality content. It checks things like document length, repetition patterns and punctuation distribution.
What NVIDIA says (2)
“Heuristic filtering employs rule-based metrics and statistical measures to identify and remove low-quality content.”
“such as document length, repetition patterns, punctuation distribution”
Large language models (LLMs) are evaluated on unseen test data. Decontamination deals with test data leaking into the training data, which would make evaluation results unreliable.
What NVIDIA says (2)
“Downstream task decontamination is a step that addresses the potential leakage of test data into training datasets”
“After training, LLMs are usually evaluated by their performance on downstream tasks consisting of unseen test data.”
Personally identifiable information (PII) ranges from direct identifiers like names and social security numbers to indirect ones. NVIDIA says PII redaction removes it to protect privacy and comply with data-protection regulations.
What NVIDIA says (2)
“Personally Identifiable Information (PII) redaction involves identifying and removing sensitive information from datasets to protect individual privacy and ensure compliance with data protection regulations.”
“from direct identifiers like names and social security numbers to indirect identifiers”
Key terms: Token Tokenization Byte Pair Encoding NeMo Curator Deduplication (exact / fuzzy / semantic) Personally identifiable information Decontamination
Try it: Tokenizer lab Deduplication lab
3.3 Testing LLM applications
How to test latency, behavior and safety of LLM apps.
Key points
Intertoken latency is the average time between consecutive generated tokens. NVIDIA notes tools differ in details; for example GenAI-Perf excludes time to first token from this average while LLMPerf includes it.
What NVIDIA says (2)
“For example, GenAI-Perf does not include TTFT in the average calculation (as opposed to LLMPerf, which does include the TTFT).”
“is the average time between the generation of consecutive tokens in a sequence. It is also known as time per output token (TPOT).”
Red teaming means assessing a system the way an attacker would, to find and reduce risks. NVIDIA's AI red team combines offensive-security professionals and data scientists to assess machine learning (ML) systems and help mitigate risks.
What NVIDIA says (2)
“Our AI red team is a cross-functional team made up of offensive security professionals and data scientists. We use our combined skills to assess our ML systems to identify and help mitigate any risks”
“mitigate any risks from the perspective of information security.”
Key terms: Time to first token Intertoken latency Red teaming
3.4 Evaluating technologies
Trading off cost, accuracy and effort between techniques.
Key points
Catastrophic forgetting is when a model loses earlier knowledge while learning new data. NVIDIA says low-rank adaptation (LoRA) reduces computational and memory cost and avoids catastrophic forgetting.
What NVIDIA says (2)
“It reduces the computational and memory cost”
“Finally, it avoids catastrophic forgetting, the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”
NVIDIA ranks customization by data and compute. Prompt engineering is light. Prompt learning needs more. Parameter-efficient fine-tuning (PEFT) needs more again. Full fine-tuning updates the pretrained weights and needs the most.
What NVIDIA says (4)
“This means fine-tuning also requires the most amount of training data and compute”
“It is light in terms of data and compute requirements.”
“providing higher accuracy than prompt engineering and prompt learning, while requiring more training data and compute.”
“This process requires more data and compute but provides better accuracy than prompt engineering.”
Quantization converts weights and activations from floating point to lower precision, typically 8-bit integers. NVIDIA says post-training quantization (PTQ) is more popular because it is simple and skips the training pipeline. Quantization-aware training (QAT) almost always gives better accuracy, so use it when PTQ accuracy is not acceptable.
What NVIDIA says (3)
“Sometimes PTQ is not able to achieve acceptable task accuracy. This is when you might consider using QAT.”
“PTQ is the more popular method of the two because it is simple and doesn’t involve the training pipeline, which also makes it the faster method. However, QAT almost always produces better accuracy”
“are converted from a floating-point representation to a lower-precision representation, typically using 8-bit integers.”
NVIDIA's retriever-evaluation post says to check the license and terms of use of pretrained models. Some pretraining datasets have licenses that prohibit commercial use.
What NVIDIA says (2)
“Another important aspect to consider is the license and terms of use associated with pretrained models”
“reviewing the datasets used for pretraining is essential as some datasets might have licenses that prohibit commercial”
Key terms: Prompt tuning / p-tuning Parameter-efficient fine-tuning Low-Rank Adaptation Catastrophic forgetting Quantization PTQ / QAT
Try it: LoRA lab
3.5 Human feedback data (RLHF)
How human rankings are collected and used to align models.
Key points
RLHF (reinforcement learning from human feedback) trains a reward model in stage 2. It uses prompts with multiple responses ranked by humans, so the reward model learns to predict human preference.
What NVIDIA says (2)
“A dataset consisting of prompts with multiple responses ranked by humans is used to train the RM to predict human preference.”
“The SFT model is trained as a reward model (RM) in stage 2 of RLHF.”
NVIDIA describes reinforcement learning from human feedback (RLHF) stage 3 as fine-tuning the initial policy model against the reward model using reinforcement learning with proximal policy optimization (PPO).
What NVIDIA says (1)
“stage 3 of RLHF focuses on fine-tuning the initial policy model against the RM using reinforcement learning with a proximal policy optimization (PPO) algorithm.”
Key terms: Reinforcement learning from human feedback Reward model Proximal policy optimization
3.6 Evaluating and benchmarking models
Choosing metrics and data that give a fair picture.
Key points
Using a large language model (LLM) to judge another LLM is common. NVIDIA warns it can introduce biases that skew results.
What NVIDIA says (1)
“Using LLMs to assess other LLMs can introduce biases that skew results, potentially compromising the accuracy of assessments.”
Recall measures how many of the relevant items were retrieved. NVIDIA recommends recall when a context of up to 4K tokens is sufficient, because NDCG (normalized discounted cumulative gain) penalizes cases where the most relevant chunk is not ranked first.
What NVIDIA says (2)
“In most information retrieval scenarios, recall is an excellent metric when the order of the retrieved candidates”
“recall is the recommended metric because NDCG penalizes when the most relevant chunk isn’t ranked at the top.”
NVIDIA says the best data to evaluate retrieval is your own. Ideally, build a clean, labeled evaluation set that reflects what you see in production.
What NVIDIA says (2)
“The best data to evaluate retrieval is your own.”
“Ideally, you build a clean and labeled evaluation dataset that best reflects what you see in production.”
Key terms: Recall / NDCG LLM-as-a-judge
3.7 Running experiments on models and pipelines
Designing experiments that tell you what to change.
Key points
NVIDIA's chunking research tested strategies across several datasets. Even within the same document category, the best strategy varied significantly. So test chunking on your own content.
What NVIDIA says (2)
“Our research evaluated different chunking strategies across multiple datasets to establish guidelines for selecting the optimal approach based on your specific content and use case.”
“Even within the same document category, optimal chunking strategies varied significantly.”
Standardized benchmarks give consistent datasets and metrics. NVIDIA warns that academic benchmarks can become saturated quickly as large language models (LLMs) improve, and tailored benchmarks for specific domains are often missing.
What NVIDIA says (3)
“Note that as LLMs develop quickly, academic benchmarks can become saturated quick”
“The lack of tailored benchmarks for specific domains limits the relevance and depth of evaluations”
“Standardized benchmarks provide consistent datasets and metrics to evaluate LLMs across a variety of tasks.”