3.1 Training and training optimization

NCA-GENL · Experimentation (22% of the exam) · Official objective: “Assist in model training and training optimization under the supervision of a senior team member.”

Techniques that make training faster or cheaper.

Key points

  1. Mixed precision runs operations in FP16 (half precision) while keeping minimal information in FP32 (single precision). NVIDIA's training curves show mixed precision without loss scaling diverging, while with loss scaling it matches the single-precision model.

    What NVIDIA says (2)

    “Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”

    — Train With Mixed Precision

    “by performing operations in half-precision format, while storing minimal information in single-precision”

    — Train With Mixed Precision

  2. Low-rank adaptation (LoRA) belongs to the parameter-efficient fine-tuning (PEFT) family. It inserts small low-rank matrices into each layer and trains only those, keeping the original large language model (LLM) weights frozen.

    What NVIDIA says (2)

    “LoRA is a fine-tuning method that introduces low-rank matrices into each layer of the LLM architecture, and only trains these matrices while keeping the original LLM weights frozen.”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “LoRA tuning is a type of tuning family called Parameter Efficient Fine-Tuning (PEFT).”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

  3. Transfer learning starts from an existing trained model. NVIDIA's example deletes the final “loss output” layer, adds a new one for the new task, then trains on the smaller new dataset: the whole network, the last few layers, or just that layer.

    What NVIDIA says (2)

    “Next, you would take your smaller dataset for horses and train it on the entire 50-layer neural network or the last few layers or just the loss layer alone.”

    — What Is Transfer Learning?

    “First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”

    — What Is Transfer Learning?

  4. The rank r is a low-rank adaptation (LoRA) hyperparameter that controls the rank of the decomposition. NVIDIA says a smaller r saves parameters and memory and trains faster, but can capture less task-specific information.

    What NVIDIA says (2)

    “Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “is a hyperparameter that controls the rank of the decomposition”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

Key terms

Try it

Sample question

In mixed-precision training, why is loss scaling used?

Show the answer

Answer: To stop small FP16 gradient values from being lost, so training matches single-precision results

Mixed precision runs operations in FP16 (half precision) while keeping minimal information in FP32 (single precision). NVIDIA's training curves show mixed precision without loss scaling diverging, while with loss scaling it matches the single-precision model.

What NVIDIA says (2)

“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”

— Train With Mixed Precision

“by performing operations in half-precision format, while storing minimal information in single-precision”

— Train With Mixed Precision

Practice 3.1 (4 questions) Full Experimentation guide

← 2.7 Writing components and scripts · 3.2 Preparing large datasets →