6.3 Transfer learning for performance
PEFT adapters, replacing the output layer, TAO export and DreamBooth.
Key points
PEFT adapts a big model by training a small number of new parameters. This cuts memory and storage compared with full fine-tuning.
What NVIDIA says (1)
“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”
The final layer maps features to the old classes. A new final layer maps them to the new classes. You then train the last layers or the whole network on the smaller dataset.
What NVIDIA says (1)
“First, you delete what’s known as the “loss output” layer, which is the final layer used to make predictions, and replace it with a new loss output layer for horse prediction.”
ONNX is an open model format that many runtimes, such as TensorRT, can load. ONNX means Open Neural Network Exchange.
What NVIDIA says (1)
“TAO outputs trained models in ONNX format”
DreamBooth is transfer learning for diffusion. It fine-tunes a pretrained model on a handful of photos of one subject.
What NVIDIA says (1)
“you only need a few images of a specific subject to fine-tune a pretrained text-to-image model”
Key terms
- DreamBooth: A fine-tuning method that teaches a text-to-image model a specific subject from a few images, bound to a unique identifier.
- Transfer learning: Reusing a model trained on one task as the start for another task.
- Parameter-efficient fine-tuning: Adapting a large model by training small added modules while the base model stays frozen.
- NVIDIA TAO: An NVIDIA toolkit for fine-tuning pretrained vision models on your data and exporting them to ONNX.
Try it
Sample question
How does NeMo 2.0 implement parameter-efficient fine-tuning (PEFT)?
Show the answer
Answer: It freezes the base model and inserts trainable adapters
PEFT adapts a big model by training a small number of new parameters. This cuts memory and storage compared with full fine-tuning.
What NVIDIA says (1)
“The new design formulates PEFT as a Model Transform that freezes the base model and inserts trainable adapters at specific locations within the model.”
Practice 6.3 (4 questions) Full Performance Optimization guide