3.3 Fixing underfitting and overfitting

NCA-ADS · Data Science Pipelines and Workflow Automation (13% of the exam) · Official objective: “Mitigating underfitting and overfitting through model and feature adjustments”

Bagging versus boosting, regularization and simpler models.

Key points

  1. Overfitting means the model learned noise in the training data. Regularization penalizes complexity so the model generalizes better.

    What NVIDIA says (2)

    “Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”

    — Gradient Boosting, Decision Trees and XGBoost with CUDA

    “Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”

    — Gradient Boosting, Decision Trees and XGBoost with CUDA

  2. Variance error comes from a model that changes too much with the data. Bias error comes from a model that is too simple. GBDT means gradient-boosted decision trees.

    What NVIDIA says (1)

    “Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”

    — NVIDIA Glossary: Random Forest

  3. Regularization adds a penalty for large weights. Ridge uses an L2 penalty; Lasso uses an L1 penalty that can set weights to zero.

    What NVIDIA says (2)

    “Ridge extends LinearRegression by providing L2 regularization on the coefficients when predicting response y with a linear combination of the predictors in X.”

    — cuML API: Ridge

    “Larger values specify stronger regularization.”

    — cuML API: Ridge

Key terms

Sample question

A gradient-boosted model scores very well on training data but poorly on new data. What is happening, and what helps?

Show the answer

Answer: Overfitting; use the objective's regularization terms or limit tree growth

Overfitting means the model learned noise in the training data. Regularization penalizes complexity so the model generalizes better.

What NVIDIA says (2)

“Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”

— Gradient Boosting, Decision Trees and XGBoost with CUDA

“Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”

— Gradient Boosting, Decision Trees and XGBoost with CUDA

Practice 3.3 (3 questions) Full Data Science Pipelines and Workflow Automation guide

← 3.2 Feature selection and transformation · 3.4 Augmenting and integrating datasets →