1.5 Testing models for accuracy
FID-CLIP, retrieval metrics, perplexity, BLEU, LLM-as-a-judge, cross-validation and robust error metrics.
Key points
FID (Fréchet Inception Distance) measures how close generated images are to real ones. A CLIP score measures how well images match their prompts. NVIDIA found both UNets similar on FID-CLIP, with slightly better visual quality from the Regular UNet. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”
Retrieval metrics check whether the right items come back. Recall@K counts how many of the relevant items appear in the top K. Precision@K counts how many of the top K are relevant. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“Assesses the proportion of relevant documents that are successfully retrieved in a grouping of K retrieved documents.”
LLM-as-a-judge uses one LLM to grade another model's output. It helps where simple metrics fall short. It can inherit bias from the judge model. LLM means large language model.
What NVIDIA says (2)
“This method excels for tasks where automated metrics fall short, such as assessing coherence and creativity.”
“it’s important to note that LLM-as-a-judge evaluations may introduce biases inherent in the LLM evaluator training data.”
Cross-validation trains and validates several times on different splits. This gives a steadier estimate of accuracy than one split.
What NVIDIA says (1)
“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”
MAE averages the absolute errors. Because it does not square them, large errors do not dominate.
What NVIDIA says (1)
“All errors are treated equally, so the metric is robust to outliers.”
Precision@K looks at what came back in the top K. Recall@K looks at how many relevant items were found. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“Measures the proportion of retrieved documents that are relevant in a grouping of K retrieved documents.”
Perplexity measures how uncertain a model is when predicting the next words. Lower uncertainty means better prediction.
What NVIDIA says (1)
“Quantifies uncertainty in predicting sequences of words, with lower values indicating better predictive performance.”
BLEU (Bilingual Evaluation Understudy) compares model output with reference translations. MAE means mean absolute error.
What NVIDIA says (2)
“Evaluates machine translation quality by comparing model outputs to reference translations.”
“The BLEU score ranges from 0 (no match, that is, low quality) to 1 (a perfect match, that is, high quality).”
Key terms
- Recall@K: The share of all relevant items that appear in the top K results.
- Precision@K: The share of the top K results that are relevant.
- FID-CLIP evaluation: Image-generation scoring that pairs FID, a realism measure, with a CLIP score for prompt match.
- LLM-as-a-judge: Using one large language model to grade another model's outputs.
- Cross-validation: Training and validating several times on different splits of the data to estimate accuracy fairly.
- Mean absolute error: The average size of prediction errors, which treats all errors equally.
- Perplexity: A measure of how uncertain a language model is when predicting text; lower is better.
Try it
Sample question
NVIDIA compared Imagen's Regular UNet and Efficient UNet. Which metric did it use for quality?
Show the answer
Answer: FID-CLIP evaluation
FID (Fréchet Inception Distance) measures how close generated images are to real ones. A CLIP score measures how well images match their prompts. NVIDIA found both UNets similar on FID-CLIP, with slightly better visual quality from the Regular UNet. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”
Practice 1.5 (8 questions) Full Experimentation guide
← 1.4 Testing multimodal data quality · 2.1 Training stability →