5.6 Parameters, tuning and fitting

NCA-ADS · Foundations of Accelerated Data Science (12% of the exam) · Official objective: “Model parameters, tuning, and overfitting vs. underfitting concepts”

Parameters versus hyperparameters, underfitting and overfitting.

Key points

  1. Training adjusts parameters, such as tree splits or linear weights. Tuning searches hyperparameters by training many times.

    What NVIDIA says (1)

    “This is particularly important because data scientists typically run the algorithm not just once, but many times in order to tune hyperparameters (such as learning rate or tree depth) and find the best accuracy.”

    — Gradient Boosting, Decision Trees and XGBoost with CUDA

  2. Underfitting is high bias: the model misses real structure. Overfitting is the opposite: it learns noise and fails on new data.

    What NVIDIA says (1)

    “Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”

    — What Is XGBoost and Why Does It Matter?

  3. Each tree overfits in its own way. Averaging their votes cancels much of that noise.

    What NVIDIA says (1)

    “The presence of a large number of trees also reduces the problem of overfitting, which occurs when a model incorporates too much “noise” in the training data and makes poor decisions as a result.”

    — NVIDIA Glossary: Random Forest

Key terms

Sample question

What is the difference between a model parameter and a hyperparameter?

Show the answer

Answer: Parameters are learned from data during training; hyperparameters such as learning rate or tree depth are set before training and tuned

Training adjusts parameters, such as tree splits or linear weights. Tuning searches hyperparameters by training many times.

What NVIDIA says (1)

“This is particularly important because data scientists typically run the algorithm not just once, but many times in order to tune hyperparameters (such as learning rate or tree depth) and find the best accuracy.”

— Gradient Boosting, Decision Trees and XGBoost with CUDA

Practice 5.6 (3 questions) Full Foundations of Accelerated Data Science guide

← 5.5 Distributed versus GPU-accelerated frameworks · 6.1 Monitoring and optimizing ML pipelines →