NCA-GENL hands-on labs

Small interactive labs that run entirely in your browser with toy data. Nothing is sent to a server and no AI model is called. Each lab is tagged to official objectives. Open one to start.

Tokenizer lab: watch BPE merges

Objectives 3.2

Byte Pair Encoding (BPE) starts from single characters and repeatedly merges the most frequent adjacent pair. Paste text and step through the merges to see the vocabulary grow and the token count shrink.

Try this

  • Run 10 merges on the sample text and note which pairs merge first.
  • Add a rare word and check how it is split into known subwords.
What NVIDIA says (2)

“BPE starts with character vocabulary and iteratively merges frequent adjacent character pairs into new vocabulary terms”

— Mastering LLM Techniques: Training

“The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”

— Mastering LLM Techniques: Training

Sampling lab: temperature, top-k and top-p

Objectives 1.9

An LLM picks each next token from a probability distribution. Change temperature, top-k and top-p on a toy distribution and see which tokens stay eligible.

Try this

  • Set temperature low and watch probability concentrate on the top token.
  • Set top-k to 1: only the most probable token remains.
  • Set top-p to 0.9 and count how many tokens are kept.
What NVIDIA says (3)

“At a lower temperature, the model is more conservative and is limited to choosing tokens with higher probabilities.”

— How to Get Better Outputs from Your Large Language Model

“When set to 1, it is always going to select the most probable token next.”

— How to Get Better Outputs from Your Large Language Model

“the model picks at random from the highest probability tokens whose probabilities sum to or exceed the top-p value.”

— How to Get Better Outputs from Your Large Language Model

Chunking lab: split a document

Objectives 1.4

Split a document into fixed-size chunks with overlap and see how chunk size changes what a retriever can return. Sizes here are counted in words as a simple stand-in for tokens.

Try this

  • Try a very small and a very large chunk size and compare the chunks.
  • Add 15% overlap and see sentences repeated across chunk boundaries.
What NVIDIA says (2)

“A chunking strategy is the method of breaking down large documents into smaller, manageable pieces for AI retrieval.”

— Finding the Best Chunking Strategy for Accurate AI Responses

“this result aligns with the 10–20% overlap commonly seen in industry practices.”

— Finding the Best Chunking Strategy for Accurate AI Responses

RAG lab: retrieve, then build the prompt

Objectives 1.3, 1.4, 1.8

A mini retrieval-augmented generation (RAG) pipeline that runs in your browser. Documents are chunked and turned into simple word-count vectors (a stand-in for a real embedding model), the query is compared by cosine similarity, and the top-K chunks are placed into a prompt. No model is called.

Try this

  • Ask a question and check that the top chunk really answers it.
  • Change K and see how the prompt grows; remember it must fit the context window.
What NVIDIA says (3)

“The system identifies relevant information by comparing the query vector with the stored vectors in the vector DBs.”

— RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

“Cosine similarity: Focuses on the angle between vectors.”

— What is a Vector Database and How Does it Work?

“The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”

— Enhancing RAG Pipelines with Re-Ranking

LoRA lab: count trainable parameters

Objectives 3.1, 3.4

Low-Rank Adaptation (LoRA) replaces training a big d × k weight matrix with two small matrices, d × r and r × k. Enter sizes and a rank r to compare trainable parameters.

Try this

  • Reproduce NVIDIA's example: 1,024 × 50,000.
  • Double r and see how trainable parameters change.
What NVIDIA says (2)

“output weight matrix \(W\) would have 1024 x 50,000 = 51,200,000 parameters.”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training.”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

KV cache lab: estimate inference memory

Objectives 1.1, 2.1

The KV cache stores attention keys and values for every token. Use NVIDIA's formula to estimate its size from batch size, sequence length, layers and hidden size.

Try this

  • Reproduce NVIDIA's example: batch 1, 4,096 tokens, 32 layers, hidden size 4,096, FP16 (2 bytes).
  • Double the batch size and see memory double.
What NVIDIA says (2)

“Total size of KV cache in bytes = (batch_size) * (sequence_length) * 2 * (num_layers) * (hidden_size) * sizeof(FP16)”

— Mastering LLM Techniques: Inference Optimization

“Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”

— Mastering LLM Techniques: Inference Optimization

Metrics lab: R², MSE, RMSE and MAE

Objectives 4.2

Paste actual and predicted values to compute regression metrics and see how one outlier moves each of them.

Try this

  • Add one large error and compare how much MSE and MAE change.
  • Predict the mean for every row and check that R² is 0.
What NVIDIA says (2)

“As the residuals are squared, MSE puts a significantly heavier penalty on large errors.”

— A Comprehensive Overview of Regression Evaluation Metrics

“In that case, RSS would be equal to TSS and result in the minimum value of R² being 0.”

— A Comprehensive Overview of Regression Evaluation Metrics

Deduplication lab: exact vs fuzzy

Objectives 3.2

Exact deduplication hashes whole documents. Fuzzy deduplication compares overlapping word shingles with Jaccard similarity to catch near-duplicates (production systems use MinHash and LSH to do this at scale).

Try this

  • Paste two documents that differ by one word: exact says different, fuzzy says similar.
  • Raise the threshold until they stop matching.
What NVIDIA says (2)

“This method generates hash signatures for each document and groups documents by their hashes into buckets”

— Mastering LLM Techniques: Text Data Processing

“Compute Jaccard similarity between documents within the same buckets.”

— Mastering LLM Techniques: Text Data Processing