Experimentation

22% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI

3.1 Training and training optimization

Official objective: “Assist in model training and training optimization under the supervision of a senior team member.”

Techniques that make training faster or cheaper.

Key points

  1. Mixed precision runs operations in FP16 (half precision) while keeping minimal information in FP32 (single precision). NVIDIA's training curves show mixed precision without loss scaling diverging, while with loss scaling it matches the single-precision model.

    What NVIDIA says (2)

    “Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”

    — Train With Mixed Precision

    “by performing operations in half-precision format, while storing minimal information in single-precision”

    — Train With Mixed Precision

  2. Low-rank adaptation (LoRA) belongs to the parameter-efficient fine-tuning (PEFT) family. It inserts small low-rank matrices into each layer and trains only those, keeping the original large language model (LLM) weights frozen.

    What NVIDIA says (2)

    “LoRA is a fine-tuning method that introduces low-rank matrices into each layer of the LLM architecture, and only trains these matrices while keeping the original LLM weights frozen.”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “LoRA tuning is a type of tuning family called Parameter Efficient Fine-Tuning (PEFT).”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

  3. Transfer learning starts from an existing trained model. NVIDIA's example deletes the final “loss output” layer, adds a new one for the new task, then trains on the smaller new dataset: the whole network, the last few layers, or just that layer.

    What NVIDIA says (2)

    “Next, you would take your smaller dataset for horses and train it on the entire 50-layer neural network or the last few layers or just the loss layer alone.”

    — What Is Transfer Learning?

    “First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”

    — What Is Transfer Learning?

  4. The rank r is a low-rank adaptation (LoRA) hyperparameter that controls the rank of the decomposition. NVIDIA says a smaller r saves parameters and memory and trains faster, but can capture less task-specific information.

    What NVIDIA says (2)

    “Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “is a hyperparameter that controls the rank of the decomposition”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

Key terms: Parameter-efficient fine-tuning Low-Rank Adaptation Supervised fine-tuning Mixed-precision training FP16 / FP32 Loss scaling Transfer learning

Try it: LoRA lab

Practice 3.1 (4 questions)

3.2 Preparing large datasets

Official objective: “Assist in preparing (e.g., scraping, tokenization) large datasets for pretraining, fine-tuning, and RLHF.”

Cleaning, deduplicating and tokenizing data for pretraining and fine-tuning.

Key points

  1. Tokenization splits text into smaller units called tokens. NVIDIA explains that word tokenizers create huge vocabularies and out-of-vocabulary tokens, while character tokenizers create long sequences. Subword tokenizers split rare words into meaningful subwords. Popular ones are byte-pair encoding (BPE), WordPiece, Unigram and SentencePiece.

    What NVIDIA says (5)

    “Word-based tokenizers lead to a large vocabulary size and words not seen during the tokenizer training process cause many out-of-vocabulary tokens.”

    — Mastering LLM Techniques: Training

    “Popular subword tokenization algorithms include Byte Pair Encoding (BPE), WordPiece, Unigram, and SentencePiece.”

    — Mastering LLM Techniques: Training

    “Character-based tokenizers lead to long sequences”

    — Mastering LLM Techniques: Training

    “The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”

    — Mastering LLM Techniques: Training

    “the process of tokenization assumes a pivotal role in fragmenting text into smaller units known as tokens.”

    — Mastering LLM Techniques: Training

  2. Exact deduplication hashes each document and keeps one per hash bucket, so it only removes identical copies. Fuzzy deduplication finds near-duplicates using MinHash signatures and Locality-Sensitive Hashing (LSH).

    What NVIDIA says (2)

    “Fuzzy deduplication addresses near-duplicate content using MinHash signatures and Locality-Sensitive Hashing (LSH) to identify similar documents.”

    — Mastering LLM Techniques: Text Data Processing

    “This method generates hash signatures for each document and groups documents by their hashes into buckets”

    — Mastering LLM Techniques: Text Data Processing

  3. NVIDIA's data-processing guide says deduplication improves training efficiency, cuts compute cost and keeps data diverse. It also helps prevent overfitting to repeated content.

    What NVIDIA says (2)

    “Deduplication is essential for improving model training efficiency, reducing computational costs, and ensuring data diversity.”

    — Mastering LLM Techniques: Text Data Processing

    “It helps prevent models from overfitting to repeated content and improves generalization.”

    — Mastering LLM Techniques: Text Data Processing

  4. NVIDIA lists Unicode fixing and language identification as crucial early steps when curating large web-scraped text corpora.

    What NVIDIA says (1)

    “Unicode fixing and language identification represent crucial early steps in the data curation pipeline, particularly when dealing with large-scale web-scraped text corpora.”

    — Mastering LLM Techniques: Text Data Processing

  5. Heuristic filtering uses rule-based metrics and statistics to remove low-quality content. It checks things like document length, repetition patterns and punctuation distribution.

    What NVIDIA says (2)

    “Heuristic filtering employs rule-based metrics and statistical measures to identify and remove low-quality content.”

    — Mastering LLM Techniques: Text Data Processing

    “such as document length, repetition patterns, punctuation distribution”

    — Mastering LLM Techniques: Text Data Processing

  6. Large language models (LLMs) are evaluated on unseen test data. Decontamination deals with test data leaking into the training data, which would make evaluation results unreliable.

    What NVIDIA says (2)

    “Downstream task decontamination is a step that addresses the potential leakage of test data into training datasets”

    — Mastering LLM Techniques: Text Data Processing

    “After training, LLMs are usually evaluated by their performance on downstream tasks consisting of unseen test data.”

    — Mastering LLM Techniques: Text Data Processing

  7. Personally identifiable information (PII) ranges from direct identifiers like names and social security numbers to indirect ones. NVIDIA says PII redaction removes it to protect privacy and comply with data-protection regulations.

    What NVIDIA says (2)

    “Personally Identifiable Information (PII) redaction involves identifying and removing sensitive information from datasets to protect individual privacy and ensure compliance with data protection regulations.”

    — Mastering LLM Techniques: Text Data Processing

    “from direct identifiers like names and social security numbers to indirect identifiers”

    — Mastering LLM Techniques: Text Data Processing

Key terms: Token Tokenization Byte Pair Encoding NeMo Curator Deduplication (exact / fuzzy / semantic) Personally identifiable information Decontamination

Try it: Tokenizer lab Deduplication lab

Practice 3.2 (7 questions)

3.3 Testing LLM applications

Official objective: “Assist in the design and conduct of hardware or software tests for LLM applications.”

How to test latency, behavior and safety of LLM apps.

Key points

  1. Intertoken latency is the average time between consecutive generated tokens. NVIDIA notes tools differ in details; for example GenAI-Perf excludes time to first token from this average while LLMPerf includes it.

    What NVIDIA says (2)

    “For example, GenAI-Perf does not include TTFT in the average calculation (as opposed to LLMPerf, which does include the TTFT).”

    — LLM Inference Benchmarking: Fundamental Concepts

    “is the average time between the generation of consecutive tokens in a sequence. It is also known as time per output token (TPOT).”

    — LLM Inference Benchmarking: Fundamental Concepts

  2. Red teaming means assessing a system the way an attacker would, to find and reduce risks. NVIDIA's AI red team combines offensive-security professionals and data scientists to assess machine learning (ML) systems and help mitigate risks.

    What NVIDIA says (2)

    “Our AI red team is a cross-functional team made up of offensive security professionals and data scientists. We use our combined skills to assess our ML systems to identify and help mitigate any risks”

    — NVIDIA AI Red Team: An Introduction

    “mitigate any risks from the perspective of information security.”

    — NVIDIA AI Red Team: An Introduction

Key terms: Time to first token Intertoken latency Red teaming

Practice 3.3 (2 questions)

3.4 Evaluating technologies

Official objective: “Assist in the evaluation of current or emerging technologies to consider factors such as cost, portability, compatibility, or usability.”

Trading off cost, accuracy and effort between techniques.

Key points

  1. Catastrophic forgetting is when a model loses earlier knowledge while learning new data. NVIDIA says low-rank adaptation (LoRA) reduces computational and memory cost and avoids catastrophic forgetting.

    What NVIDIA says (2)

    “It reduces the computational and memory cost”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “Finally, it avoids catastrophic forgetting, the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

  2. NVIDIA ranks customization by data and compute. Prompt engineering is light. Prompt learning needs more. Parameter-efficient fine-tuning (PEFT) needs more again. Full fine-tuning updates the pretrained weights and needs the most.

    What NVIDIA says (4)

    “This means fine-tuning also requires the most amount of training data and compute”

    — Mastering LLM Techniques: Customization

    “It is light in terms of data and compute requirements.”

    — Mastering LLM Techniques: Customization

    “providing higher accuracy than prompt engineering and prompt learning, while requiring more training data and compute.”

    — Mastering LLM Techniques: Customization

    “This process requires more data and compute but provides better accuracy than prompt engineering.”

    — Mastering LLM Techniques: Customization

  3. Quantization converts weights and activations from floating point to lower precision, typically 8-bit integers. NVIDIA says post-training quantization (PTQ) is more popular because it is simple and skips the training pipeline. Quantization-aware training (QAT) almost always gives better accuracy, so use it when PTQ accuracy is not acceptable.

    What NVIDIA says (3)

    “Sometimes PTQ is not able to achieve acceptable task accuracy. This is when you might consider using QAT.”

    — Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

    “PTQ is the more popular method of the two because it is simple and doesn’t involve the training pipeline, which also makes it the faster method. However, QAT almost always produces better accuracy”

    — Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

    “are converted from a floating-point representation to a lower-precision representation, typically using 8-bit integers.”

    — Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

  4. NVIDIA's retriever-evaluation post says to check the license and terms of use of pretrained models. Some pretraining datasets have licenses that prohibit commercial use.

    What NVIDIA says (2)

    “Another important aspect to consider is the license and terms of use associated with pretrained models”

    — Evaluating Retriever for Enterprise-Grade RAG

    “reviewing the datasets used for pretraining is essential as some datasets might have licenses that prohibit commercial”

    — Evaluating Retriever for Enterprise-Grade RAG

Key terms: Prompt tuning / p-tuning Parameter-efficient fine-tuning Low-Rank Adaptation Catastrophic forgetting Quantization PTQ / QAT

Try it: LoRA lab

Practice 3.4 (4 questions)

3.5 Human feedback data (RLHF)

Official objective: “Awareness of, and/or participation in, data collection from human subjects (e.g., RLHF).”

How human rankings are collected and used to align models.

Key points

  1. RLHF (reinforcement learning from human feedback) trains a reward model in stage 2. It uses prompts with multiple responses ranked by humans, so the reward model learns to predict human preference.

    What NVIDIA says (2)

    “A dataset consisting of prompts with multiple responses ranked by humans is used to train the RM to predict human preference.”

    — Mastering LLM Techniques: Customization

    “The SFT model is trained as a reward model (RM) in stage 2 of RLHF.”

    — Mastering LLM Techniques: Customization

  2. NVIDIA describes reinforcement learning from human feedback (RLHF) stage 3 as fine-tuning the initial policy model against the reward model using reinforcement learning with proximal policy optimization (PPO).

    What NVIDIA says (1)

    “stage 3 of RLHF focuses on fine-tuning the initial policy model against the RM using reinforcement learning with a proximal policy optimization (PPO) algorithm.”

    — Mastering LLM Techniques: Customization

Key terms: Reinforcement learning from human feedback Reward model Proximal policy optimization

Practice 3.5 (2 questions)

3.6 Evaluating and benchmarking models

Official objective: “Evaluate and refine existing models / benchmarking.”

Choosing metrics and data that give a fair picture.

Key points

  1. Using a large language model (LLM) to judge another LLM is common. NVIDIA warns it can introduce biases that skew results.

    What NVIDIA says (1)

    “Using LLMs to assess other LLMs can introduce biases that skew results, potentially compromising the accuracy of assessments.”

    — Mastering LLM Techniques: Evaluation

  2. Recall measures how many of the relevant items were retrieved. NVIDIA recommends recall when a context of up to 4K tokens is sufficient, because NDCG (normalized discounted cumulative gain) penalizes cases where the most relevant chunk is not ranked first.

    What NVIDIA says (2)

    “In most information retrieval scenarios, recall is an excellent metric when the order of the retrieved candidates”

    — Evaluating Retriever for Enterprise-Grade RAG

    “recall is the recommended metric because NDCG penalizes when the most relevant chunk isn’t ranked at the top.”

    — Evaluating Retriever for Enterprise-Grade RAG

  3. NVIDIA says the best data to evaluate retrieval is your own. Ideally, build a clean, labeled evaluation set that reflects what you see in production.

    What NVIDIA says (2)

    “The best data to evaluate retrieval is your own.”

    — Evaluating Retriever for Enterprise-Grade RAG

    “Ideally, you build a clean and labeled evaluation dataset that best reflects what you see in production.”

    — Evaluating Retriever for Enterprise-Grade RAG

Key terms: Recall / NDCG LLM-as-a-judge

Practice 3.6 (3 questions)

3.7 Running experiments on models and pipelines

Official objective: “Executes experimentation to evaluate models and pipelines.”

Designing experiments that tell you what to change.

Key points

  1. NVIDIA's chunking research tested strategies across several datasets. Even within the same document category, the best strategy varied significantly. So test chunking on your own content.

    What NVIDIA says (2)

    “Our research evaluated different chunking strategies across multiple datasets to establish guidelines for selecting the optimal approach based on your specific content and use case.”

    — Finding the Best Chunking Strategy for Accurate AI Responses

    “Even within the same document category, optimal chunking strategies varied significantly.”

    — Finding the Best Chunking Strategy for Accurate AI Responses

  2. Standardized benchmarks give consistent datasets and metrics. NVIDIA warns that academic benchmarks can become saturated quickly as large language models (LLMs) improve, and tailored benchmarks for specific domains are often missing.

    What NVIDIA says (3)

    “Note that as LLMs develop quickly, academic benchmarks can become saturated quick”

    — Mastering LLM Techniques: Evaluation

    “The lack of tailored benchmarks for specific domains limits the relevance and depth of evaluations”

    — Mastering LLM Techniques: Evaluation

    “Standardized benchmarks provide consistent datasets and metrics to evaluate LLMs across a variety of tasks.”

    — Mastering LLM Techniques: Evaluation

Practice 3.7 (2 questions)