3.3 Fixing underfitting and overfitting
Bagging versus boosting, regularization and simpler models.
Key points
Overfitting means the model learned noise in the training data. Regularization penalizes complexity so the model generalizes better.
What NVIDIA says (2)
“Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”
“Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”
Variance error comes from a model that changes too much with the data. Bias error comes from a model that is too simple. GBDT means gradient-boosted decision trees.
What NVIDIA says (1)
“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”
Regularization adds a penalty for large weights. Ridge uses an L2 penalty; Lasso uses an L1 penalty that can set weights to zero.
What NVIDIA says (2)
“Ridge extends LinearRegression by providing L2 regularization on the coefficients when predicting response y with a linear combination of the predictors in X.”
“Larger values specify stronger regularization.”
Key terms
- Random forest: Many decision trees trained on random samples whose votes are combined.
- Bagging: Training many models on random samples and averaging them to reduce variance.
- Boosting: Training models one after another, each fixing the errors of the last, to reduce bias.
- Regularization: A penalty on model complexity that helps prevent overfitting.
- Overfitting: When a model learns noise in the training data and does poorly on new data.
- Underfitting: When a model is too simple to capture the pattern, even on training data.
Sample question
A gradient-boosted model scores very well on training data but poorly on new data. What is happening, and what helps?
Show the answer
Answer: Overfitting; use the objective's regularization terms or limit tree growth
Overfitting means the model learned noise in the training data. Regularization penalizes complexity so the model generalizes better.
What NVIDIA says (2)
“Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”
“Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”
Practice 3.3 (3 questions) Full Data Science Pipelines and Workflow Automation guide
← 3.2 Feature selection and transformation · 3.4 Augmenting and integrating datasets →