1.5 Testing models for accuracy

NCA-GENM · Experimentation (25% of the exam) · Official objective: “Test AI models to ensure their accuracy and effectiveness.”

FID-CLIP, retrieval metrics, perplexity, BLEU, LLM-as-a-judge, cross-validation and robust error metrics.

Key points

  1. FID (Fréchet Inception Distance) measures how close generated images are to real ones. A CLIP score measures how well images match their prompts. NVIDIA found both UNets similar on FID-CLIP, with slightly better visual quality from the Regular UNet. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”

    — NeMo Framework 24.09: Imagen

  2. Retrieval metrics check whether the right items come back. Recall@K counts how many of the relevant items appear in the top K. Precision@K counts how many of the top K are relevant. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “Assesses the proportion of relevant documents that are successfully retrieved in a grouping of K retrieved documents.”

    — Mastering LLM Techniques: Evaluation

  3. LLM-as-a-judge uses one LLM to grade another model's output. It helps where simple metrics fall short. It can inherit bias from the judge model. LLM means large language model.

    What NVIDIA says (2)

    “This method excels for tasks where automated metrics fall short, such as assessing coherence and creativity.”

    — Mastering LLM Techniques: Evaluation

    “it’s important to note that LLM-as-a-judge evaluations may introduce biases inherent in the LLM evaluator training data.”

    — Mastering LLM Techniques: Evaluation

  4. Cross-validation trains and validates several times on different splits. This gives a steadier estimate of accuracy than one split.

    What NVIDIA says (1)

    “Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”

    — What is scikit-learn?

  5. MAE averages the absolute errors. Because it does not square them, large errors do not dominate.

    What NVIDIA says (1)

    “All errors are treated equally, so the metric is robust to outliers.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  6. Precision@K looks at what came back in the top K. Recall@K looks at how many relevant items were found. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “Measures the proportion of retrieved documents that are relevant in a grouping of K retrieved documents.”

    — Mastering LLM Techniques: Evaluation

  7. Perplexity measures how uncertain a model is when predicting the next words. Lower uncertainty means better prediction.

    What NVIDIA says (1)

    “Quantifies uncertainty in predicting sequences of words, with lower values indicating better predictive performance.”

    — Mastering LLM Techniques: Evaluation

  8. BLEU (Bilingual Evaluation Understudy) compares model output with reference translations. MAE means mean absolute error.

    What NVIDIA says (2)

    “Evaluates machine translation quality by comparing model outputs to reference translations.”

    — Mastering LLM Techniques: Evaluation

    “The BLEU score ranges from 0 (no match, that is, low quality) to 1 (a perfect match, that is, high quality).”

    — Mastering LLM Techniques: Evaluation

Key terms

Try it

Sample question

NVIDIA compared Imagen's Regular UNet and Efficient UNet. Which metric did it use for quality?

Show the answer

Answer: FID-CLIP evaluation

FID (Fréchet Inception Distance) measures how close generated images are to real ones. A CLIP score measures how well images match their prompts. NVIDIA found both UNets similar on FID-CLIP, with slightly better visual quality from the Regular UNet. CLIP means Contrastive Language-Image Pre-training.

What NVIDIA says (1)

“Metric-wise, they exhibit similar quality based on FID-CLIP evaluation.”

— NeMo Framework 24.09: Imagen

Practice 1.5 (8 questions) Full Experimentation guide

← 1.4 Testing multimodal data quality · 2.1 Training stability →