6.3 Transfer learning for performance

NCA-GENM · Performance Optimization (10% of the exam) · Official objective: “Develop content for multimodal-specific transfer learning.”

PEFT adapters, replacing the output layer, TAO export and DreamBooth.

Key points

  1. PEFT adapts a big model by training a small number of new parameters. This cuts memory and storage compared with full fine-tuning.

    What NVIDIA says (1)

    “The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”

    — NeMo Framework 24.09: Parameter-Efficient Fine-Tuning (NeMo 2.0)

  2. The final layer maps features to the old classes. A new final layer maps them to the new classes. You then train the last layers or the whole network on the smaller dataset.

    What NVIDIA says (1)

    “First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”

    — What Is Transfer Learning?

  3. ONNX is an open model format that many runtimes, such as TensorRT, can load. ONNX means Open Neural Network Exchange.

    What NVIDIA says (1)

    “TAO outputs trained models in ONNX format”

    — NVIDIA TAO Toolkit: Overview

  4. DreamBooth is transfer learning for diffusion. It fine-tunes a pretrained model on a handful of photos of one subject.

    What NVIDIA says (1)

    “you only need a few images of a specific subject to fine-tune a pretrained text-to-image model”

    — NeMo Framework 24.09: DreamBooth

Key terms

Try it

Sample question

How does NeMo 2.0 implement parameter-efficient fine-tuning (PEFT)?

Show the answer

Answer: It freezes the base model and inserts trainable adapters

PEFT adapts a big model by training a small number of new parameters. This cuts memory and storage compared with full fine-tuning.

What NVIDIA says (1)

“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”

— NeMo Framework 24.09: Parameter-Efficient Fine-Tuning (NeMo 2.0)

Practice 6.3 (4 questions) Full Performance Optimization guide

← 6.2 Hyperparameter tuning · 6.4 Training optimization →