2.5 Statistics for evaluating pipelines

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Design statistical analyses for evaluating multimodal pipelines.”

What R-squared, RMSE and ROUGE tell you, and what they do not.

Key points

  1. R-squared is the share of variance in the target that a model explains. A high value alone does not show the model is unbiased or general.

    What NVIDIA says (1)

    “R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”

    — A Comprehensive Overview of Regression Evaluation Metrics

  2. RMSE (root mean squared error) is on the target's scale. But it is not the average error, because errors are squared first.

    What NVIDIA says (1)

    “an RMSE of 10 does not actually mean you are off by 10 units on average.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  3. ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is a text overlap metric. Higher ROUGE means the summary shares more with the reference.

    What NVIDIA says (2)

    “Measures overlap between machine-generated and human-generated summaries.”

    — Mastering LLM Techniques: Evaluation

    “The ROUGE score ranges between 0 and 1, with higher scores indicating higher similarity.”

    — Mastering LLM Techniques: Evaluation

Key terms

Try it

Sample question

A model scores a high R-squared. Why can that still mislead you?

Show the answer

Answer: R-squared gives no measure of bias, so an overfitted model can still score high

R-squared is the share of variance in the target that a model explains. A high value alone does not show the model is unbiased or general.

What NVIDIA says (1)

“R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”

— A Comprehensive Overview of Regression Evaluation Metrics

Practice 2.5 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 2.4 Nonsequential networks and residual connections · 2.6 Multimodal transfer learning →