3.1 Training and training optimization
Techniques that make training faster or cheaper.
Key points
Mixed precision runs operations in FP16 (half precision) while keeping minimal information in FP32 (single precision). NVIDIA's training curves show mixed precision without loss scaling diverging, while with loss scaling it matches the single-precision model.
What NVIDIA says (2)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
“by performing operations in half-precision format, while storing minimal information in single-precision”
Low-rank adaptation (LoRA) belongs to the parameter-efficient fine-tuning (PEFT) family. It inserts small low-rank matrices into each layer and trains only those, keeping the original large language model (LLM) weights frozen.
What NVIDIA says (2)
“LoRA is a fine-tuning method that introduces low-rank matrices into each layer of the LLM architecture, and only trains these matrices while keeping the original LLM weights frozen.”
“LoRA tuning is a type of tuning family called Parameter Efficient Fine-Tuning (PEFT).”
Transfer learning starts from an existing trained model. NVIDIA's example deletes the final “loss output” layer, adds a new one for the new task, then trains on the smaller new dataset: the whole network, the last few layers, or just that layer.
What NVIDIA says (2)
“Next, you would take your smaller dataset for horses and train it on the entire 50-layer neural network or the last few layers or just the loss layer alone.”
“First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”
The rank r is a low-rank adaptation (LoRA) hyperparameter that controls the rank of the decomposition. NVIDIA says a smaller r saves parameters and memory and trains faster, but can capture less task-specific information.
What NVIDIA says (2)
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”
“is a hyperparameter that controls the rank of the decomposition”
Key terms
- Parameter-efficient fine-tuning: Fine-tuning that adds or updates only a few parameters or layers instead of the whole model.
- Low-Rank Adaptation: A PEFT method that trains small low-rank matrices in each layer while the original weights stay frozen.
- Supervised fine-tuning: Fine-tuning the model's weights on labeled examples, such as instructions with desired answers.
- Mixed-precision training: Training that does most math in 16-bit floats while keeping key values in 32-bit.
- FP16 / FP32: 16-bit and 32-bit floating-point number formats.
- Loss scaling: Scaling the loss in mixed-precision training so small gradients are not lost.
- Transfer learning: Starting from a model trained on one task and adapting it to a related task.
Try it
Sample question
In mixed-precision training, why is loss scaling used?
Show the answer
Answer: To stop small FP16 gradient values from being lost, so training matches single-precision results
Mixed precision runs operations in FP16 (half precision) while keeping minimal information in FP32 (single precision). NVIDIA's training curves show mixed precision without loss scaling diverging, while with loss scaling it matches the single-precision model.
What NVIDIA says (2)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
“by performing operations in half-precision format, while storing minimal information in single-precision”
Practice 3.1 (4 questions) Full Experimentation guide
← 2.7 Writing components and scripts · 3.2 Preparing large datasets →