NCA-GENL hands-on labs
Small interactive labs that run entirely in your browser with toy data. Nothing is sent to a server and no AI model is called. Each lab is tagged to official objectives. Open one to start.
Tokenizer lab: watch BPE merges
Byte Pair Encoding (BPE) starts from single characters and repeatedly merges the most frequent adjacent pair. Paste text and step through the merges to see the vocabulary grow and the token count shrink.
Try this
- Run 10 merges on the sample text and note which pairs merge first.
- Add a rare word and check how it is split into known subwords.
What NVIDIA says (2)
“BPE starts with character vocabulary and iteratively merges frequent adjacent character pairs into new vocabulary terms”
“The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”
Sampling lab: temperature, top-k and top-p
An LLM picks each next token from a probability distribution. Change temperature, top-k and top-p on a toy distribution and see which tokens stay eligible.
Try this
- Set temperature low and watch probability concentrate on the top token.
- Set top-k to 1: only the most probable token remains.
- Set top-p to 0.9 and count how many tokens are kept.
What NVIDIA says (3)
“At a lower temperature, the model is more conservative and is limited to choosing tokens with higher probabilities.”
“When set to 1, it is always going to select the most probable token next.”
“the model picks at random from the highest probability tokens whose probabilities sum to or exceed the top-p value.”
Chunking lab: split a document
Split a document into fixed-size chunks with overlap and see how chunk size changes what a retriever can return. Sizes here are counted in words as a simple stand-in for tokens.
Try this
- Try a very small and a very large chunk size and compare the chunks.
- Add 15% overlap and see sentences repeated across chunk boundaries.
What NVIDIA says (2)
“A chunking strategy is the method of breaking down large documents into smaller, manageable pieces for AI retrieval.”
“this result aligns with the 10–20% overlap commonly seen in industry practices.”
RAG lab: retrieve, then build the prompt
A mini retrieval-augmented generation (RAG) pipeline that runs in your browser. Documents are chunked and turned into simple word-count vectors (a stand-in for a real embedding model), the query is compared by cosine similarity, and the top-K chunks are placed into a prompt. No model is called.
Try this
- Ask a question and check that the top chunk really answers it.
- Change K and see how the prompt grows; remember it must fit the context window.
What NVIDIA says (3)
“The system identifies relevant information by comparing the query vector with the stored vectors in the vector DBs.”
“Cosine similarity: Focuses on the angle between vectors.”
“The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”
LoRA lab: count trainable parameters
Low-Rank Adaptation (LoRA) replaces training a big d × k weight matrix with two small matrices, d × r and r × k. Enter sizes and a rank r to compare trainable parameters.
Try this
- Reproduce NVIDIA's example: 1,024 × 50,000.
- Double r and see how trainable parameters change.
What NVIDIA says (2)
“output weight matrix \(W\) would have 1024 x 50,000 = 51,200,000 parameters.”
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training.”
KV cache lab: estimate inference memory
The KV cache stores attention keys and values for every token. Use NVIDIA's formula to estimate its size from batch size, sequence length, layers and hidden size.
Try this
- Reproduce NVIDIA's example: batch 1, 4,096 tokens, 32 layers, hidden size 4,096, FP16 (2 bytes).
- Double the batch size and see memory double.
What NVIDIA says (2)
“Total size of KV cache in bytes = (batch_size) * (sequence_length) * 2 * (num_layers) * (hidden_size) * sizeof(FP16)”
“Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”
Metrics lab: R², MSE, RMSE and MAE
Paste actual and predicted values to compute regression metrics and see how one outlier moves each of them.
Try this
- Add one large error and compare how much MSE and MAE change.
- Predict the mean for every row and check that R² is 0.
What NVIDIA says (2)
“As the residuals are squared, MSE puts a significantly heavier penalty on large errors.”
“In that case, RSS would be equal to TSS and result in the minimum value of R² being 0.”
Deduplication lab: exact vs fuzzy
Exact deduplication hashes whole documents. Fuzzy deduplication compares overlapping word shingles with Jaccard similarity to catch near-duplicates (production systems use MinHash and LSH to do this at scale).
Try this
- Paste two documents that differ by one word: exact says different, fuzzy says similar.
- Raise the threshold until they stop matching.
What NVIDIA says (2)
“This method generates hash signatures for each document and groups documents by their hashes into buckets”
“Compute Jaccard similarity between documents within the same buckets.”