3.4 Evaluating technologies
Trading off cost, accuracy and effort between techniques.
Key points
Catastrophic forgetting is when a model loses earlier knowledge while learning new data. NVIDIA says low-rank adaptation (LoRA) reduces computational and memory cost and avoids catastrophic forgetting.
What NVIDIA says (2)
“It reduces the computational and memory cost”
“Finally, it avoids catastrophic forgetting, the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”
NVIDIA ranks customization by data and compute. Prompt engineering is light. Prompt learning needs more. Parameter-efficient fine-tuning (PEFT) needs more again. Full fine-tuning updates the pretrained weights and needs the most.
What NVIDIA says (4)
“This means fine-tuning also requires the most amount of training data and compute”
“It is light in terms of data and compute requirements.”
“providing higher accuracy than prompt engineering and prompt learning, while requiring more training data and compute.”
“This process requires more data and compute but provides better accuracy than prompt engineering.”
Quantization converts weights and activations from floating point to lower precision, typically 8-bit integers. NVIDIA says post-training quantization (PTQ) is more popular because it is simple and skips the training pipeline. Quantization-aware training (QAT) almost always gives better accuracy, so use it when PTQ accuracy is not acceptable.
What NVIDIA says (3)
“Sometimes PTQ is not able to achieve acceptable task accuracy. This is when you might consider using QAT.”
“PTQ is the more popular method of the two because it is simple and doesn’t involve the training pipeline, which also makes it the faster method. However, QAT almost always produces better accuracy”
“are converted from a floating-point representation to a lower-precision representation, typically using 8-bit integers.”
NVIDIA's retriever-evaluation post says to check the license and terms of use of pretrained models. Some pretraining datasets have licenses that prohibit commercial use.
What NVIDIA says (2)
“Another important aspect to consider is the license and terms of use associated with pretrained models”
“reviewing the datasets used for pretraining is essential as some datasets might have licenses that prohibit commercial”
Key terms
- Prompt tuning / p-tuning: Learning small virtual-token embeddings for a task while the LLM's own weights stay frozen.
- Parameter-efficient fine-tuning: Fine-tuning that adds or updates only a few parameters or layers instead of the whole model.
- Low-Rank Adaptation: A PEFT method that trains small low-rank matrices in each layer while the original weights stay frozen.
- Catastrophic forgetting: When a model loses earlier knowledge while learning new data.
- Quantization: Storing weights and activations in a lower-precision format, typically 8-bit integers.
- PTQ / QAT: PTQ quantizes a trained model; QAT includes quantization during training for better accuracy.
Try it
Sample question
Compared with full fine-tuning, why might a team choose a parameter-efficient method like low-rank adaptation (LoRA)?
Show the answer
Answer: It reduces computational and memory cost and avoids catastrophic forgetting
Catastrophic forgetting is when a model loses earlier knowledge while learning new data. NVIDIA says low-rank adaptation (LoRA) reduces computational and memory cost and avoids catastrophic forgetting.
What NVIDIA says (2)
“It reduces the computational and memory cost”
“Finally, it avoids catastrophic forgetting, the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”
Practice 3.4 (4 questions) Full Experimentation guide
← 3.3 Testing LLM applications · 3.5 Human feedback data (RLHF) →