3.6 Evaluating and benchmarking models

NCA-GENL · Experimentation (22% of the exam) · Official objective: “Evaluate and refine existing models / benchmarking.”

Choosing metrics and data that give a fair picture.

Key points

  1. Using a large language model (LLM) to judge another LLM is common. NVIDIA warns it can introduce biases that skew results.

    What NVIDIA says (1)

    “Using LLMs to assess other LLMs can introduce biases that skew results, potentially compromising the accuracy of assessments.”

    — Mastering LLM Techniques: Evaluation

  2. Recall measures how many of the relevant items were retrieved. NVIDIA recommends recall when a context of up to 4K tokens is sufficient, because NDCG (normalized discounted cumulative gain) penalizes cases where the most relevant chunk is not ranked first.

    What NVIDIA says (2)

    “In most information retrieval scenarios, recall is an excellent metric when the order of the retrieved candidates”

    — Evaluating Retriever for Enterprise-Grade RAG

    “recall is the recommended metric because NDCG penalizes when the most relevant chunk isn’t ranked at the top.”

    — Evaluating Retriever for Enterprise-Grade RAG

  3. NVIDIA says the best data to evaluate retrieval is your own. Ideally, build a clean, labeled evaluation set that reflects what you see in production.

    What NVIDIA says (2)

    “The best data to evaluate retrieval is your own.”

    — Evaluating Retriever for Enterprise-Grade RAG

    “Ideally, you build a clean and labeled evaluation dataset that best reflects what you see in production.”

    — Evaluating Retriever for Enterprise-Grade RAG

Key terms

Sample question

A team uses a large language model (LLM) to grade another LLM's answers. What risk does NVIDIA warn about?

Show the answer

Answer: The judge can introduce biases that skew results

Using a large language model (LLM) to judge another LLM is common. NVIDIA warns it can introduce biases that skew results.

What NVIDIA says (1)

“Using LLMs to assess other LLMs can introduce biases that skew results, potentially compromising the accuracy of assessments.”

— Mastering LLM Techniques: Evaluation

Practice 3.6 (3 questions) Full Experimentation guide

← 3.5 Human feedback data (RLHF) · 3.7 Running experiments on models and pipelines →