3.4 Evaluating technologies

NCA-GENL · Experimentation (22% of the exam) · Official objective: “Assist in the evaluation of current or emerging technologies to consider factors such as cost, portability, compatibility, or usability.”

Trading off cost, accuracy and effort between techniques.

Key points

  1. Catastrophic forgetting is when a model loses earlier knowledge while learning new data. NVIDIA says low-rank adaptation (LoRA) reduces computational and memory cost and avoids catastrophic forgetting.

    What NVIDIA says (2)

    “It reduces the computational and memory cost”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “Finally, it avoids catastrophic forgetting, the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

  2. NVIDIA ranks customization by data and compute. Prompt engineering is light. Prompt learning needs more. Parameter-efficient fine-tuning (PEFT) needs more again. Full fine-tuning updates the pretrained weights and needs the most.

    What NVIDIA says (4)

    “This means fine-tuning also requires the most amount of training data and compute”

    — Mastering LLM Techniques: Customization

    “It is light in terms of data and compute requirements.”

    — Mastering LLM Techniques: Customization

    “providing higher accuracy than prompt engineering and prompt learning, while requiring more training data and compute.”

    — Mastering LLM Techniques: Customization

    “This process requires more data and compute but provides better accuracy than prompt engineering.”

    — Mastering LLM Techniques: Customization

  3. Quantization converts weights and activations from floating point to lower precision, typically 8-bit integers. NVIDIA says post-training quantization (PTQ) is more popular because it is simple and skips the training pipeline. Quantization-aware training (QAT) almost always gives better accuracy, so use it when PTQ accuracy is not acceptable.

    What NVIDIA says (3)

    “Sometimes PTQ is not able to achieve acceptable task accuracy. This is when you might consider using QAT.”

    — Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

    “PTQ is the more popular method of the two because it is simple and doesn’t involve the training pipeline, which also makes it the faster method. However, QAT almost always produces better accuracy”

    — Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

    “are converted from a floating-point representation to a lower-precision representation, typically using 8-bit integers.”

    — Achieving FP32 Accuracy for INT8 Inference Using Quantization Aware Training with NVIDIA TensorRT

  4. NVIDIA's retriever-evaluation post says to check the license and terms of use of pretrained models. Some pretraining datasets have licenses that prohibit commercial use.

    What NVIDIA says (2)

    “Another important aspect to consider is the license and terms of use associated with pretrained models”

    — Evaluating Retriever for Enterprise-Grade RAG

    “reviewing the datasets used for pretraining is essential as some datasets might have licenses that prohibit commercial”

    — Evaluating Retriever for Enterprise-Grade RAG

Key terms

Try it

Sample question

Compared with full fine-tuning, why might a team choose a parameter-efficient method like low-rank adaptation (LoRA)?

Show the answer

Answer: It reduces computational and memory cost and avoids catastrophic forgetting

Catastrophic forgetting is when a model loses earlier knowledge while learning new data. NVIDIA says low-rank adaptation (LoRA) reduces computational and memory cost and avoids catastrophic forgetting.

What NVIDIA says (2)

“It reduces the computational and memory cost”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

“Finally, it avoids catastrophic forgetting, the natural tendency of LLMs to abruptly forget previously learned information upon learning new data.”

— Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

Practice 3.4 (4 questions) Full Experimentation guide

← 3.3 Testing LLM applications · 3.5 Human feedback data (RLHF) →