2.3 Evaluating and comparing models

NCA-ADS · Machine Learning With RAPIDS (16% of the exam) · Official objective: “Model evaluation, comparison, and generalization assessment”

Held-out test sets, R-squared, and judging generalization.

Key points

  1. Generalization is how well a model does on data it has not seen. A held-out test set measures it; training accuracy alone can hide overfitting.

    What NVIDIA says (2)

    “test_size float or int, default=None If float, should be between 0.0 and 1.0 and represent the proportion of the dataset to include”

    — cuML API: train_test_split

    “Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  2. R², the coefficient of determination, is the share of the target's variance that a model explains. It is a relative metric for models trained on the same data.

    What NVIDIA says (2)

    “R-squared (R²), also known as the coefficient of determination, represents the proportion of variance explained by a model.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “R² is a relative metric; that is, it can be used to compare with other models trained on the same dataset.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  3. A model that always predicts the mean gets R² = 0. A negative score means the model is worse than that simple baseline.

    What NVIDIA says (1)

    “Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”

    — cuML API: r2_score

  4. Pick the metric that fits the task. Accuracy is the ratio of correct predictions to all predictions; R² fits continuous targets.

    What NVIDIA says (2)

    “Accuracy score is the ratio of correct predictions to the total number of predictions.”

    — cuML: Training and evaluating machine learning models

    “The cell below uses the Linear Regression model and evaluates its performance using cuML’s R² score metric.”

    — cuML: Training and evaluating machine learning models

Key terms

Try it

Sample question

Why hold out a test set that the model never sees during training?

Show the answer

Answer: To estimate how well the model generalizes to new data

Generalization is how well a model does on data it has not seen. A held-out test set measures it; training accuracy alone can hide overfitting.

What NVIDIA says (2)

“test_size float or int, default=None If float, should be between 0.0 and 1.0 and represent the proportion of the dataset to include”

— cuML API: train_test_split

“Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

Practice 2.3 (4 questions) Full Machine Learning With RAPIDS guide

← 2.2 Regression, classification and clustering · 2.4 Hyperparameter tuning →