Performance Optimization
10% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI
6.1 Efficiency versus accuracy
Tensor Core alignment, memory-bound layers, BF16 and activation recomputation.
Key points
Tensor Cores are GPU units for fast matrix math. They work best when sizes line up with their tile shapes: multiples of 4 for TF32, 8 for FP16 and 16 for INT8. FP16 means 16-bit floating point. TF32 means TensorFloat-32, a 19-bit Tensor Core math format. INT8 means 8-bit integer.
What NVIDIA says (1)
“Tensor Cores are most efficient when key parameters of the operation are multiples of 4 if using TF32, 8 if using FP16, or 16 if using INT8”
Math-limited means compute speed is the bottleneck. Memory-bound means moving data is the bottleneck. Small layer parameters often make an operation memory-bound.
What NVIDIA says (1)
“speeding up calculation does not improve performance.”
BF16 is a 16-bit format with the same exponent range as FP32. It saves memory and time while staying numerically stable. CLIP means Contrastive Language-Image Pre-training. FP16 means 16-bit floating point. FP32 means 32-bit floating point.
What NVIDIA says (1)
“Training is conducted in Bfloat16 precision, which offers a balance between the higher precision of FP32 and the memory savings and speed of FP16.”
Activation recomputation stores only some activations and recomputes the rest in the backward pass. It trades extra compute for less memory.
What NVIDIA says (2)
“Checkpointing a few activations and recomputing the rest is a common technique to reduce device memory usage.”
“it increases the per-transformer layer computation cost by 30%”
Key terms: Mixed precision bfloat16 Tensor Core Memory-bound Activation recomputation
6.2 Hyperparameter tuning
Optimizer and schedule configs, Hydra overrides, LoRA rank and batch sizing.
Key points
Learning rate and warmup are hyperparameters, set before training. Warmup ramps the learning rate up gradually at the start.
What NVIDIA says (1)
“Optimizers and learning rate schedules are configurable across all NeMo models and have their own namespace.”
Hydra merges settings from several layers. Command-line overrides are handy for quick tuning sweeps. YAML means a plain-text configuration file format.
What NVIDIA says (1)
“Configuration with Hydra always has the following precedence CLI > YAML > Dataclass.”
LoRA (Low-Rank Adaptation) trains small low-rank matrices while freezing the base model. The rank r is a hyperparameter you tune.
What NVIDIA says (2)
“is a hyperparameter that controls the rank of the decomposition”
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”
GPUs split work into tiles. Sizes that divide evenly avoid wasted work.
What NVIDIA says (1)
“choosing parameters (including batch size, input size, output size, and channel counts) to be divisible by larger powers of two, at least 64, and up to 256.”
Key terms: Low-Rank Adaptation Hydra Hyperparameter
Try it: LoRA lab
6.3 Transfer learning for performance
PEFT adapters, replacing the output layer, TAO export and DreamBooth.
Key points
PEFT adapts a big model by training a small number of new parameters. This cuts memory and storage compared with full fine-tuning.
What NVIDIA says (1)
“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”
The final layer maps features to the old classes. A new final layer maps them to the new classes. You then train the last layers or the whole network on the smaller dataset.
What NVIDIA says (1)
“First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”
ONNX is an open model format that many runtimes, such as TensorRT, can load. ONNX means Open Neural Network Exchange.
What NVIDIA says (1)
“TAO outputs trained models in ONNX format”
DreamBooth is transfer learning for diffusion. It fine-tunes a pretrained model on a handful of photos of one subject.
What NVIDIA says (1)
“you only need a few images of a specific subject to fine-tune a pretrained text-to-image model”
Key terms: DreamBooth Transfer learning Parameter-efficient fine-tuning NVIDIA TAO
Try it: LoRA lab
6.4 Training optimization
Sequence packing, data parallelism and precomputed encodings.
Key points
Padding fills short sequences up to a fixed length, wasting compute. Sequence packing joins several examples into one long sequence instead. LLM means large language model.
What NVIDIA says (2)
“Many sequences are short, and a few are very long, conforming to Zipf’s Law.”
“Sequence packing is a training technique wherein multiple training sequences (examples) are concatenated into one long sequence (pack).”
Each GPU processes its share of the batch. Gradients are then synchronized so the copies stay the same.
What NVIDIA says (1)
“Data Parallelism (DP) replicates the model across multiple GPUs. Data batches are evenly distributed between GPUs and the data-parallel GPUs process them independently.”
Frozen encoders always give the same output, so you can compute it once and reuse it. VAE means variational autoencoder.
What NVIDIA says (1)
“Since the VAE and text encoder remain frozen during training, you can pre-calculate the image and caption latents offline, enhancing training throughput.”
Key terms: Sequence packing Data parallelism