NCA-GENM hands-on labs

Small interactive labs that run entirely in your browser with toy data. Nothing is sent to a server and no AI model is called. Each lab is tagged to official objectives. Open one to start.

Diffusion lab: from pure noise to a clean signal

Objectives 4.4

Denoising diffusion starts from pure random noise and removes a little noise at each step. This lab runs NVIDIA's toy sampler loop on an 8-value signal and shows the noise level and the distance to the clean signal at each step.

Try this

  • Run 256 steps at 0.98 and check how far the result is from the clean signal.
  • Cut the steps to 20 and see how much noise is left.
What NVIDIA says (2)

“first draw a random image of pure white noise, and then chip away at the noise level”

— NVIDIA Technical Blog: Demystifying Diffusion-Based Models

“the denoiser must output the blurry average of all possible clean images that could have been hiding under the noise.”

— NVIDIA Technical Blog: Demystifying Diffusion-Based Models

CLIP lab: match images and captions

Objectives 1.1, 2.2, 4.5

CLIP scores how well each image matches each caption. Contrastive training raises the score of the matching pair and lowers the rest. Here, short tag lists stand in for images and word-count vectors stand in for the encoders, so you can see the image-by-caption matrix.

Try this

  • Check that each image picks the caption on its own row.
  • Swap two captions and watch the matches move off the diagonal.
What NVIDIA says (2)

“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

— NeMo Framework 24.09: CLIP

“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

— NeMo Framework 24.09: CLIP

Sampling lab: temperature, top-k and top-p

Objectives 2.9, 4.3

A generative model picks each next token from a probability distribution. Change temperature, top-k and top-p on a toy distribution and see which tokens stay eligible.

Try this

  • Set temperature low and watch probability concentrate on the top token.
  • Set top-k to 1: only the most probable token remains.
What NVIDIA says (2)

“Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”

— How to Get Better Outputs from Your Large Language Model

“Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”

— How to Get Better Outputs from Your Large Language Model

RAG lab: retrieve, then build the prompt

Objectives 3.2, 3.3

A mini retrieval-augmented generation (RAG) pipeline. Documents are chunked and turned into word-count vectors (a stand-in for a real embedding model). The query vector is compared with each chunk by cosine similarity, and the top K chunks go into the prompt.

Try this

  • Ask a question and check that the top chunk really answers it.
  • Change K and see how the prompt grows.
What NVIDIA says (2)

“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”

— RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

“refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”

— What is a Vector Database and How Does it Work?

LoRA lab: count trainable parameters

Objectives 6.2, 6.3

Low-Rank Adaptation (LoRA) trains two small matrices instead of a big weight matrix while the base model stays frozen. Enter sizes and a rank r to compare trainable parameters.

Try this

  • Halve r and see how many parameters you save.
  • Compare r = 4 with r = 64 for the same layer.
What NVIDIA says (2)

“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”

— NeMo Framework 24.09: Parameter-Efficient Fine-Tuning (NeMo 2.0)

Metrics lab: R², MSE, RMSE and MAE

Objectives 1.5, 2.5

Paste actual and predicted values to compute regression metrics and see how one outlier moves each of them.

Try this

  • Add one large error and compare how much RMSE and MAE change.
  • Predict the mean for every row and check that R² is 0.
What NVIDIA says (2)

“All errors are treated equally, so the metric is robust to outliers.”

— A Comprehensive Overview of Regression Evaluation Metrics

“R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”

— A Comprehensive Overview of Regression Evaluation Metrics

Deduplication lab: exact vs near-duplicate

Objectives 1.4

Exact deduplication only catches identical items. Near-duplicates need a similarity score and a threshold. This lab uses word shingles and Jaccard similarity; NeMo Curator's semantic deduplication applies the same threshold idea to embeddings with cosine similarity.

Try this

  • Paste two captions that differ by one word: exact says different, the similarity score says near-duplicate.
  • Raise the threshold until they stop matching.
What NVIDIA says (2)

“uses embeddings to identify and remove “semantic duplicates” - data pairs that are semantically similar but not exactly identical.”

— NeMo Curator (NeMo Framework 24.09): Semantic Deduplication

“Data pairs with cosine similarity above a threshold are considered semantic duplicates.”

— NeMo Curator (NeMo Framework 24.09): Semantic Deduplication