4.2 Comparing models with statistics

NCA-GENL · Data Analysis and Visualization (14% of the exam) · Official objective: “Compare models using statistical performance metrics, such as loss functions or proportion of explained variance.”

Regression metrics and what they do and do not tell you.

Key points

  1. R² (the coefficient of determination) is the proportion of variance in the target that the model explains. A higher value means a better fit when you compare models trained on the same dataset.

    What NVIDIA says (3)

    “also known as the coefficient of determination, represents the proportion of variance explained by a model.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “R² is a relative metric; that is, it can be used to compare with other models trained on the same dataset. A higher value indicates a better fit.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “corresponds to the degree to which the variance in the dependent variable (the target) can be explained by the”

    — A Comprehensive Overview of Regression Evaluation Metrics

  2. A residual is the difference between the actual and the predicted value. Mean squared error (MSE) is the average of the squared residuals. Squaring puts a much heavier penalty on large errors, so MSE is not robust to outliers.

    What NVIDIA says (3)

    “As the residuals are squared, MSE puts a significantly heavier penalty on large errors. Some of those might be outliers, so MSE is not robust to their presence.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “a residual is a difference between the actual value and the predicted value.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “The difference is that you are now interested in the average error instead of the total error.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  3. MAE (mean absolute error) uses absolute values, so it ignores the direction of errors. Like mean squared error (MSE) and RMSE, it is scale-dependent, so you cannot compare it between different datasets.

    What NVIDIA says (2)

    “Similar to MSE and RMSE, MAE is also scale-dependent, so you cannot compare it between different datasets.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “Absolute value disregards the direction of the errors”

    — A Comprehensive Overview of Regression Evaluation Metrics

  4. RMSE is the square root of mean squared error (MSE). Taking the root brings the metric back to the scale of the target variable, so it is easier to interpret.

    What NVIDIA says (2)

    “(RMSE) is closely related to MSE, as it is simply the square root of the latter.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “Take the square to bring the metric back to the scale of the target variable, so it is easier to interpret and understand.”

    — A Comprehensive Overview of Regression Evaluation Metrics

Key terms

Try it

Sample question

What does R-squared (R², the coefficient of determination) tell you about a regression model?

Show the answer

Answer: The proportion of variance in the target explained by the model

R² (the coefficient of determination) is the proportion of variance in the target that the model explains. A higher value means a better fit when you compare models trained on the same dataset.

What NVIDIA says (3)

“also known as the coefficient of determination, represents the proportion of variance explained by a model.”

— A Comprehensive Overview of Regression Evaluation Metrics

“R² is a relative metric; that is, it can be used to compare with other models trained on the same dataset. A higher value indicates a better fit.”

— A Comprehensive Overview of Regression Evaluation Metrics

“corresponds to the degree to which the variance in the dependent variable (the target) can be explained by the”

— A Comprehensive Overview of Regression Evaluation Metrics

Practice 4.2 (4 questions) Full Data Analysis and Visualization guide

← 4.1 Insights from large datasets · 4.3 Doing the data analysis →