NCA-GENM hands-on labs
Small interactive labs that run entirely in your browser with toy data. Nothing is sent to a server and no AI model is called. Each lab is tagged to official objectives. Open one to start.
Diffusion lab: from pure noise to a clean signal
Denoising diffusion starts from pure random noise and removes a little noise at each step. This lab runs NVIDIA's toy sampler loop on an 8-value signal and shows the noise level and the distance to the clean signal at each step.
Try this
- Run 256 steps at 0.98 and check how far the result is from the clean signal.
- Cut the steps to 20 and see how much noise is left.
What NVIDIA says (2)
“first draw a random image of pure white noise, and then chip away at the noise level”
“the denoiser must output the blurry average of all possible clean images that could have been hiding under the noise.”
CLIP lab: match images and captions
CLIP scores how well each image matches each caption. Contrastive training raises the score of the matching pair and lowers the rest. Here, short tag lists stand in for images and word-count vectors stand in for the encoders, so you can see the image-by-caption matrix.
Try this
- Check that each image picks the caption on its own row.
- Swap two captions and watch the matches move off the diagonal.
What NVIDIA says (2)
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
Sampling lab: temperature, top-k and top-p
A generative model picks each next token from a probability distribution. Change temperature, top-k and top-p on a toy distribution and see which tokens stay eligible.
Try this
- Set temperature low and watch probability concentrate on the top token.
- Set top-k to 1: only the most probable token remains.
What NVIDIA says (2)
“Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”
“Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”
RAG lab: retrieve, then build the prompt
A mini retrieval-augmented generation (RAG) pipeline. Documents are chunked and turned into word-count vectors (a stand-in for a real embedding model). The query vector is compared with each chunk by cosine similarity, and the top K chunks go into the prompt.
Try this
- Ask a question and check that the top chunk really answers it.
- Change K and see how the prompt grows.
What NVIDIA says (2)
“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
“refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
LoRA lab: count trainable parameters
Low-Rank Adaptation (LoRA) trains two small matrices instead of a big weight matrix while the base model stays frozen. Enter sizes and a rank r to compare trainable parameters.
Try this
- Halve r and see how many parameters you save.
- Compare r = 4 with r = 64 for the same layer.
What NVIDIA says (2)
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”
“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”
Metrics lab: R², MSE, RMSE and MAE
Paste actual and predicted values to compute regression metrics and see how one outlier moves each of them.
Try this
- Add one large error and compare how much RMSE and MAE change.
- Predict the mean for every row and check that R² is 0.
What NVIDIA says (2)
“All errors are treated equally, so the metric is robust to outliers.”
“R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”
Deduplication lab: exact vs near-duplicate
Exact deduplication only catches identical items. Near-duplicates need a similarity score and a threshold. This lab uses word shingles and Jaccard similarity; NeMo Curator's semantic deduplication applies the same threshold idea to embeddings with cosine similarity.
Try this
- Paste two captions that differ by one word: exact says different, the similarity score says near-duplicate.
- Raise the threshold until they stop matching.
What NVIDIA says (2)
“uses embeddings to identify and remove “semantic duplicates” - data pairs that are semantically similar but not exactly identical.”
“Data pairs with cosine similarity above a threshold are considered semantic duplicates.”