Core Machine Learning and AI Knowledge
20% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI
2.1 Training stability
Convergence settings in CLIP and loss scaling in mixed precision.
Key points
Convergence means the loss settles to a good value as training goes on. gather_with_grad keeps full gradients when features are gathered across GPUs. Turning it off can make training unstable. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Disabling this (setting to False ) may cause convergence issues.”
Mixed precision trains mostly in 16-bit floating point (FP16) for speed. Small gradients can round to zero in FP16. Loss scaling multiplies the loss so those gradients stay representable. FP32 means 32-bit floating point. INT8 means 8-bit integer.
What NVIDIA says (1)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
An overflow means a value became too large (Inf or NaN). Dynamic loss scaling then skips that step and lowers the scale. If no overflow happens for a while, it raises the scale again.
What NVIDIA says (1)
“If an overflow occurs, skip the weight update and decrease the scaling factor.”
Key terms: Loss scaling Mixed precision
2.2 Multimodal loss functions
Contrastive loss, the diffusion denoising loss and DreamBooth prior preservation.
Key points
A loss function scores how wrong a model is; training lowers it. CLIP's loss compares every image with every caption in a batch. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
A diffusion model learns to remove noise step by step. Its denoiser, often a U-Net, is trained with mean squared error (MSE). CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Training a denoiser network (typically a U-Net) with basic loss—the mean square error between its output and the clean target—achieves precisely this result.”
Language drift is when the model forgets what a general word means after narrow fine-tuning. Prior preservation loss guides the model with its own generated samples of the general class. Its weight is set with model.prior_loss_weight. VAE means variational autoencoder. KL means Kullback-Leibler (a distance between distributions). L1 means absolute-difference.
What NVIDIA says (2)
“problems like language drift and decreased output variety often arise.”
“it guides the model using its self-generated samples and incorporates the discrepancy between the model-predicted noise on these samples.”
CLIP compares every image with every caption, which builds a large matrix across GPUs. local_loss avoids building the full global matrix. This helps memory when training on many devices. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“local_loss : If set to True , the loss is calculated with local features at a global level, avoiding the need to realize the full global matrix.”
Key terms: Contrastive loss Prior preservation loss
Try it: CLIP lab
2.3 Machine learning fundamentals
Feature engineering, model comparison with grid search and cross-validation, and supervised versus unsupervised learning.
Key points
Here a transformer is a scikit-learn preprocessing step, not a neural network. Feature engineering means shaping raw data into useful inputs, such as scaled numbers.
What NVIDIA says (2)
“Transformers apply algorithms to clean or reshape the training data before it is fed into a model.”
“For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”
A hyperparameter is a setting chosen before training, such as tree depth. Grid search tries many settings. Cross-validation scores each fairly on several splits.
What NVIDIA says (1)
“For effective model selection, scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”
A label is the known answer for a training example. Clustering is a common unsupervised task.
What NVIDIA says (1)
“Supervised learning algorithms use labeled data, unsupervised learning algorithms find patterns in unlabeled data.”
Key terms: Cross-validation scikit-learn Feature engineering Hyperparameter
2.4 Nonsequential networks and residual connections
Skip connections in ResNet and U-Nets, and ControlNet side branches.
Key points
A residual connection adds a block's input to its output, skipping the layers in between. A nonsequential network has paths like this that do not run strictly one layer after another. The skip path lets gradients reach early layers.
What NVIDIA says (1)
“ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”
A U-Net has down-sampling and up-sampling paths joined by skip connections. Efficient UNet moves parameters to low-resolution blocks and scales the skips. NVIDIA says it converges faster and uses memory better.
What NVIDIA says (2)
“adding more residual blocks for the lower resolutions”
“Scaling skip connection by 1/sqrt(2)”
ControlNet is a side branch beside the original network. The locked copy preserves the pretrained model. The trainable copy learns the new condition.
What NVIDIA says (1)
“It copies the weights of neural network blocks into a “locked” copy and a “trainable” copy.”
Key terms: U-Net ControlNet Residual (skip) connection
2.5 Statistics for evaluating pipelines
What R-squared, RMSE and ROUGE tell you, and what they do not.
Key points
R-squared is the share of variance in the target that a model explains. A high value alone does not show the model is unbiased or general.
What NVIDIA says (1)
“R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”
RMSE (root mean squared error) is on the target's scale. But it is not the average error, because errors are squared first.
What NVIDIA says (1)
“an RMSE of 10 does not actually mean you are off by 10 units on average.”
ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is a text overlap metric. Higher ROUGE means the summary shares more with the reference.
What NVIDIA says (2)
“Measures overlap between machine-generated and human-generated summaries.”
“The ROUGE score ranges between 0 and 1, with higher scores indicating higher similarity.”
Key terms: Mean absolute error Root mean squared error R-squared ROUGE
Try it: Metrics lab
2.6 Multimodal transfer learning
Freezing pretrained parts, when transfer learning helps, and TAO fine-tuning.
Key points
Transfer learning reuses a pretrained model for a new task. Freezing keeps pretrained parts fixed while new parts, such as a projection layer, learn.
What NVIDIA says (1)
“freeze : If set to True , the model parameters will not be updated during training.”
Transfer learning applies a network trained on one task to another domain. It helps most when new data is scarce.
What NVIDIA says (2)
“This deep learning technique enables developers to harness a neural network used for one task and apply it to another domain.”
“Transfer learning is useful when you have insufficient data for a new domain”
TAO is NVIDIA's toolkit for adapting pretrained models. ONNX (Open Neural Network Exchange) is a standard model file format.
What NVIDIA says (2)
“You can select from 100+ pre-trained vision AI models on NGC and fine-tune them on your own dataset”
“TAO outputs trained models in ONNX format”
Key terms: DreamBooth Transfer learning Freezing NVIDIA TAO
2.7 Emerging multimodal trends
Vision language models and world foundation models such as NVIDIA Cosmos.
Key points
A VLM combines a large language model with a vision encoder so the LLM can 'see'. Unlike fixed-class vision models, it can follow natural-language instructions.
What NVIDIA says (1)
“Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”
Zero-shot means doing a task with no task-specific training examples. Traditional CV models must be retrained when a new class is added. VLMs means vision language models.
What NVIDIA says (1)
“Out of the box, VLMs have strong zero-shot performance on a variety of vision tasks”
Physical AI is AI that perceives and acts in the physical world, such as robots and vehicles. Cosmos provides world foundation models that generate and predict video of the world.
What NVIDIA says (1)
“NVIDIA Cosmos is a developer-first platform for designing Physical AI systems.”
Key terms: Multimodal model Vision language model NVIDIA Cosmos Zero-shot
2.8 Energy-efficient, trustworthy models
Quantizing the costly part, curating less data, and being transparent.
Key points
Quantization stores and computes with fewer bits, such as INT8, to save time and energy. NVIDIA uses ModelOpt to calibrate and quantize the SDXL UNet. VAE means variational autoencoder. INT8 means 8-bit integer.
What NVIDIA says (1)
“The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”
Less data means fewer training steps and less energy. Removing near-duplicates keeps most of the useful signal.
What NVIDIA says (1)
“it can remove up to 50% of the data with minimal performance loss.”
Trustworthy AI puts safety and transparency first. Being transparent includes sharing benchmarks and dataset descriptions.
What NVIDIA says (1)
“They’re also transparent — providing information such as accuracy benchmarks or a description of the training dataset”
Key terms: Semantic deduplication Quantization Trustworthy AI
2.9 Prompt engineering principles
Zero-shot and few-shot prompts, chain of thought and temperature.
Key points
Prompt engineering is writing model inputs to get the output you want. A zero-shot prompt simply asks, with no examples.
What NVIDIA says (1)
“Zero-shot means prompting the model without any example of expected behavior from the model.”
Few-shot prompting includes examples in the prompt. When the examples show reasoning, the model tends to show its own. LLM means large language model.
What NVIDIA says (1)
“Do this by providing some few-shot examples, where the reasoning process is explained. When the LLM answers the prompt, it shows its reasoning process as well.”
Temperature controls how random the token choice is. At low temperature the model picks higher-probability tokens.
What NVIDIA says (1)
“Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”
Key terms: Prompt engineering Zero-shot Few-shot prompting Temperature
Try it: Sampling lab
2.10 TensorFlow and PyTorch
Automatic mixed precision, the NeMo training stack and data loading in PyTorch.
Key points
Automatic mixed precision lets the framework pick FP16 or FP32 per operation. In supported frameworks it can take one line of code or one environment variable. FP16 means 16-bit floating point. FP32 means 32-bit floating point. NLP means natural-language processing. ML means machine learning.
What NVIDIA says (1)
“Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”
NeMo is a PyTorch-based framework. PyTorch Lightning runs the training loop. Hydra manages the YAML configs. YAML means a plain-text configuration file format.
What NVIDIA says (1)
“NeMo uses Hydra for configuring both NeMo models and the PyTorch Lightning Trainer.”
In PyTorch you can swap the data loader for DALI. The example enables it with a --data-backends flag. DALI means the NVIDIA Data Loading Library.
What NVIDIA says (1)
“DALI can use CPU or GPU, and outperforms the PyTorch native dataloader.”
Key terms: NVIDIA DALI PyTorch Lightning Hydra Automatic mixed precision