Core Machine Learning and AI Knowledge

20% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI

2.1 Training stability

Official objective: “Control training stability in multimodal settings.”

Convergence settings in CLIP and loss scaling in mixed precision.

Key points

  1. Convergence means the loss settles to a good value as training goes on. gather_with_grad keeps full gradients when features are gathered across GPUs. Turning it off can make training unstable. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Disabling this (setting to False ) may cause convergence issues.”

    — NeMo Framework 24.09: CLIP

  2. Mixed precision trains mostly in 16-bit floating point (FP16) for speed. Small gradients can round to zero in FP16. Loss scaling multiplies the loss so those gradients stay representable. FP32 means 32-bit floating point. INT8 means 8-bit integer.

    What NVIDIA says (1)

    “Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”

    — Train With Mixed Precision

  3. An overflow means a value became too large (Inf or NaN). Dynamic loss scaling then skips that step and lowers the scale. If no overflow happens for a while, it raises the scale again.

    What NVIDIA says (1)

    “If an overflow occurs, skip the weight update and decrease the scaling factor.”

    — Train With Mixed Precision

Key terms: Loss scaling Mixed precision

Practice 2.1 (3 questions)

2.2 Multimodal loss functions

Official objective: “Develop content introducing multimodal loss functions.”

Contrastive loss, the diffusion denoising loss and DreamBooth prior preservation.

Key points

  1. A loss function scores how wrong a model is; training lowers it. CLIP's loss compares every image with every caption in a batch. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

    — NeMo Framework 24.09: CLIP

  2. A diffusion model learns to remove noise step by step. Its denoiser, often a U-Net, is trained with mean squared error (MSE). CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Training a denoiser network (typically a U-Net) with basic loss—the mean square error between its output and the clean target—achieves precisely this result.”

    — NVIDIA Technical Blog: Demystifying Diffusion-Based Models

  3. Language drift is when the model forgets what a general word means after narrow fine-tuning. Prior preservation loss guides the model with its own generated samples of the general class. Its weight is set with model.prior_loss_weight. VAE means variational autoencoder. KL means Kullback-Leibler (a distance between distributions). L1 means absolute-difference.

    What NVIDIA says (2)

    “problems like language drift and decreased output variety often arise.”

    — NeMo Framework 24.09: DreamBooth

    “it guides the model using its self-generated samples and incorporates the discrepancy between the model-predicted noise on these samples.”

    — NeMo Framework 24.09: DreamBooth

  4. CLIP compares every image with every caption, which builds a large matrix across GPUs. local_loss avoids building the full global matrix. This helps memory when training on many devices. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “local_loss : If set to True , the loss is calculated with local features at a global level, avoiding the need to realize the full global matrix.”

    — NeMo Framework 24.09: CLIP

Key terms: Contrastive loss Prior preservation loss

Try it: CLIP lab

Practice 2.2 (4 questions)

2.3 Machine learning fundamentals

Official objective: “Apply machine learning fundamentals (feature engineering, model comparison, cross-validation).”

Feature engineering, model comparison with grid search and cross-validation, and supervised versus unsupervised learning.

Key points

  1. Here a transformer is a scikit-learn preprocessing step, not a neural network. Feature engineering means shaping raw data into useful inputs, such as scaled numbers.

    What NVIDIA says (2)

    “Transformers apply algorithms to clean or reshape the training data before it is fed into a model.”

    — What is scikit-learn?

    “For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”

    — What is scikit-learn?

  2. A hyperparameter is a setting chosen before training, such as tree depth. Grid search tries many settings. Cross-validation scores each fairly on several splits.

    What NVIDIA says (1)

    “For effective model selection, scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”

    — What is scikit-learn?

  3. A label is the known answer for a training example. Clustering is a common unsupervised task.

    What NVIDIA says (1)

    “Supervised learning algorithms use labeled data, unsupervised learning algorithms find patterns in unlabeled data.”

    — What is Machine Learning and Why Does It Matter?

Key terms: Cross-validation scikit-learn Feature engineering Hyperparameter

Practice 2.3 (3 questions)

2.4 Nonsequential networks and residual connections

Official objective: “Explain nonsequential neural networks and residual connections.”

Skip connections in ResNet and U-Nets, and ControlNet side branches.

Key points

  1. A residual connection adds a block's input to its output, skipping the layers in between. A nonsequential network has paths like this that do not run strictly one layer after another. The skip path lets gradients reach early layers.

    What NVIDIA says (1)

    “ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”

    — NVIDIA Technical Blog: Accelerating AI Training with MLPerf Containers and Models from NGC

  2. A U-Net has down-sampling and up-sampling paths joined by skip connections. Efficient UNet moves parameters to low-resolution blocks and scales the skips. NVIDIA says it converges faster and uses memory better.

    What NVIDIA says (2)

    “adding more residual blocks for the lower resolutions”

    — NeMo Framework 24.09: Imagen

    “Scaling skip connection by 1/sqrt(2)”

    — NeMo Framework 24.09: Imagen

  3. ControlNet is a side branch beside the original network. The locked copy preserves the pretrained model. The trainable copy learns the new condition.

    What NVIDIA says (1)

    “It copies the weights of neural network blocks into a “locked” copy and a “trainable” copy.”

    — NeMo Framework 24.09: ControlNet

Key terms: U-Net ControlNet Residual (skip) connection

Practice 2.4 (3 questions)

2.5 Statistics for evaluating pipelines

Official objective: “Design statistical analyses for evaluating multimodal pipelines.”

What R-squared, RMSE and ROUGE tell you, and what they do not.

Key points

  1. R-squared is the share of variance in the target that a model explains. A high value alone does not show the model is unbiased or general.

    What NVIDIA says (1)

    “R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”

    — A Comprehensive Overview of Regression Evaluation Metrics

  2. RMSE (root mean squared error) is on the target's scale. But it is not the average error, because errors are squared first.

    What NVIDIA says (1)

    “an RMSE of 10 does not actually mean you are off by 10 units on average.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  3. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is a text overlap metric. Higher ROUGE means the summary shares more with the reference.

    What NVIDIA says (2)

    “Measures overlap between machine-generated and human-generated summaries.”

    — Mastering LLM Techniques: Evaluation

    “The ROUGE score ranges between 0 and 1, with higher scores indicating higher similarity.”

    — Mastering LLM Techniques: Evaluation

Key terms: Mean absolute error Root mean squared error R-squared ROUGE

Try it: Metrics lab

Practice 2.5 (3 questions)

2.6 Multimodal transfer learning

Official objective: “Develop content for multimodal-specific transfer learning.”

Freezing pretrained parts, when transfer learning helps, and TAO fine-tuning.

Key points

  1. Transfer learning reuses a pretrained model for a new task. Freezing keeps pretrained parts fixed while new parts, such as a projection layer, learn.

    What NVIDIA says (1)

    “freeze : If set to True , the model parameters will not be updated during training.”

    — NeMo Framework 24.09: NeVA

  2. Transfer learning applies a network trained on one task to another domain. It helps most when new data is scarce.

    What NVIDIA says (2)

    “This deep learning technique enables developers to harness a neural network used for one task and apply it to another domain.”

    — What Is Transfer Learning?

    “Transfer learning is useful when you have insufficient data for a new domain”

    — What Is Transfer Learning?

  3. TAO is NVIDIA's toolkit for adapting pretrained models. ONNX (Open Neural Network Exchange) is a standard model file format.

    What NVIDIA says (2)

    “You can select from 100+ pre-trained vision AI models on NGC and fine-tune them on your own dataset”

    — NVIDIA TAO Toolkit: Overview

    “TAO outputs trained models in ONNX format”

    — NVIDIA TAO Toolkit: Overview

Key terms: DreamBooth Transfer learning Freezing NVIDIA TAO

Practice 2.6 (3 questions)

2.7 Emerging multimodal trends

Official objective: “Track emerging multimodal trends and technologies.”

Vision language models and world foundation models such as NVIDIA Cosmos.

Key points

  1. A VLM combines a large language model with a vision encoder so the LLM can 'see'. Unlike fixed-class vision models, it can follow natural-language instructions.

    What NVIDIA says (1)

    “Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”

    — NVIDIA Glossary: Vision Language Models

  2. Zero-shot means doing a task with no task-specific training examples. Traditional CV models must be retrained when a new class is added. VLMs means vision language models.

    What NVIDIA says (1)

    “Out of the box, VLMs have strong zero-shot performance on a variety of vision tasks”

    — NVIDIA Glossary: Vision Language Models

  3. Physical AI is AI that perceives and acts in the physical world, such as robots and vehicles. Cosmos provides world foundation models that generate and predict video of the world.

    What NVIDIA says (1)

    “NVIDIA Cosmos is a developer-first platform for designing Physical AI systems.”

    — NVIDIA Cosmos: Introduction

Key terms: Multimodal model Vision language model NVIDIA Cosmos Zero-shot

Practice 2.7 (3 questions)

2.8 Energy-efficient, trustworthy models

Official objective: “Contribute to the design, development, and deployment of energy-efficient, trustworthy multimodal AI models.”

Quantizing the costly part, curating less data, and being transparent.

Key points

  1. Quantization stores and computes with fewer bits, such as INT8, to save time and energy. NVIDIA uses ModelOpt to calibrate and quantize the SDXL UNet. VAE means variational autoencoder. INT8 means 8-bit integer.

    What NVIDIA says (1)

    “The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”

    — NeMo Framework 24.09: Stable Diffusion XL Int8 Quantization

  2. Less data means fewer training steps and less energy. Removing near-duplicates keeps most of the useful signal.

    What NVIDIA says (1)

    “it can remove up to 50% of the data with minimal performance loss.”

    — NeMo Curator (NeMo Framework 24.09): Semantic Deduplication

  3. Trustworthy AI puts safety and transparency first. Being transparent includes sharing benchmarks and dataset descriptions.

    What NVIDIA says (1)

    “They’re also transparent — providing information such as accuracy benchmarks or a description of the training dataset”

    — What Is Trustworthy AI?

Key terms: Semantic deduplication Quantization Trustworthy AI

Practice 2.8 (3 questions)

2.9 Prompt engineering principles

Official objective: “Apply prompt engineering principles to craft prompts that achieve desired results.”

Zero-shot and few-shot prompts, chain of thought and temperature.

Key points

  1. Prompt engineering is writing model inputs to get the output you want. A zero-shot prompt simply asks, with no examples.

    What NVIDIA says (1)

    “Zero-shot means prompting the model without any example of expected behavior from the model.”

    — An Introduction to Large Language Models: Prompt Engineering and P-Tuning

  2. Few-shot prompting includes examples in the prompt. When the examples show reasoning, the model tends to show its own. LLM means large language model.

    What NVIDIA says (1)

    “Do this by providing some few-shot examples, where the reasoning process is explained. When the LLM answers the prompt, it shows its reasoning process as well.”

    — An Introduction to Large Language Models: Prompt Engineering and P-Tuning

  3. Temperature controls how random the token choice is. At low temperature the model picks higher-probability tokens.

    What NVIDIA says (1)

    “Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”

    — How to Get Better Outputs from Your Large Language Model

Key terms: Prompt engineering Zero-shot Few-shot prompting Temperature

Try it: Sampling lab

Practice 2.9 (3 questions)

2.10 TensorFlow and PyTorch

Official objective: “Work with deep learning frameworks such as TensorFlow or PyTorch.”

Automatic mixed precision, the NeMo training stack and data loading in PyTorch.

Key points

  1. Automatic mixed precision lets the framework pick FP16 or FP32 per operation. In supported frameworks it can take one line of code or one environment variable. FP16 means 16-bit floating point. FP32 means 32-bit floating point. NLP means natural-language processing. ML means machine learning.

    What NVIDIA says (1)

    “Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”

    — Train With Mixed Precision

  2. NeMo is a PyTorch-based framework. PyTorch Lightning runs the training loop. Hydra manages the YAML configs. YAML means a plain-text configuration file format.

    What NVIDIA says (1)

    “NeMo uses Hydra for configuring both NeMo models and the PyTorch Lightning Trainer.”

    — NeMo Framework 24.09: NeMo Models (core)

  3. In PyTorch you can swap the data loader for DALI. The example enables it with a --data-backends flag. DALI means the NVIDIA Data Loading Library.

    What NVIDIA says (1)

    “DALI can use CPU or GPU, and outperforms the PyTorch native dataloader.”

    — NVIDIA DeepLearningExamples: ResNet50 v1.5 for PyTorch

Key terms: NVIDIA DALI PyTorch Lightning Hydra Automatic mixed precision

Practice 2.10 (3 questions)