3.6 Evaluating and benchmarking models
Choosing metrics and data that give a fair picture.
Key points
Using a large language model (LLM) to judge another LLM is common. NVIDIA warns it can introduce biases that skew results.
What NVIDIA says (1)
“Using LLMs to assess other LLMs can introduce biases that skew results, potentially compromising the accuracy of assessments.”
Recall measures how many of the relevant items were retrieved. NVIDIA recommends recall when a context of up to 4K tokens is sufficient, because NDCG (normalized discounted cumulative gain) penalizes cases where the most relevant chunk is not ranked first.
What NVIDIA says (2)
“In most information retrieval scenarios, recall is an excellent metric when the order of the retrieved candidates”
“recall is the recommended metric because NDCG penalizes when the most relevant chunk isn’t ranked at the top.”
NVIDIA says the best data to evaluate retrieval is your own. Ideally, build a clean, labeled evaluation set that reflects what you see in production.
What NVIDIA says (2)
“The best data to evaluate retrieval is your own.”
“Ideally, you build a clean and labeled evaluation dataset that best reflects what you see in production.”
Key terms
- Recall / NDCG: Retrieval metrics: recall counts relevant items found; NDCG also rewards ranking them high.
- LLM-as-a-judge: Using one LLM to grade another LLM's answers.
Sample question
A team uses a large language model (LLM) to grade another LLM's answers. What risk does NVIDIA warn about?
Show the answer
Answer: The judge can introduce biases that skew results
Using a large language model (LLM) to judge another LLM is common. NVIDIA warns it can introduce biases that skew results.
What NVIDIA says (1)
“Using LLMs to assess other LLMs can introduce biases that skew results, potentially compromising the accuracy of assessments.”
Practice 3.6 (3 questions) Full Experimentation guide
← 3.5 Human feedback data (RLHF) · 3.7 Running experiments on models and pipelines →