2.4 Hyperparameter tuning

NCA-ADS · Machine Learning With RAPIDS (16% of the exam) · Official objective: “Hyperparameter tuning and optimization”

Grid search, random search and Optuna with cuML.

Key points

  1. A hyperparameter is a setting chosen before training, unlike weights learned during training. HPO searches for good settings.

    What NVIDIA says (1)

    “Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  2. Grid search tries every combination of the values you list. That is 3 x 3 = 9 here, and the grid grows fast as you add parameters.

    What NVIDIA says (2)

    “The grid search will take place over |n_estimators| x |max_depth| which is 3 x 3 = 9.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

    “As you have probably guessed, the grid size grows rapidly as the number of parameters and their search space increases.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  3. Random search draws a fixed number of combinations at random. In NVIDIA's notebook it matched grid search with just 25 combinations.

    What NVIDIA says (2)

    “Random Search replaces the exhaustive nature of the search from before with a random selection of parameters over the specified space.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

    “We notice that performing grid search and random search yields similar performance improvements even though random search used just 25 combination of parameters.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  4. Optuna is a lightweight framework for automatic HPO. You give it an objective function that trains and scores a model; it chooses the next settings. HPO means hyperparameter optimization.

    What NVIDIA says (2)

    “Optuna is a lightweight framework for automatic hyperparameter optimization.”

    — RAPIDS Deployment: Optuna HPO with RAPIDS

    “By simply wrapping the objective function with Optuna, we can perform a parallel-distributed HPO search over a search space as we’ll see in this notebook.”

    — RAPIDS Deployment: Optuna HPO with RAPIDS

Key terms

Sample question

What is hyperparameter optimization (HPO)?

Show the answer

Answer: Choosing settings such as tree depth or number of trees that give the best results on held-out data

A hyperparameter is a setting chosen before training, unlike weights learned during training. HPO searches for good settings.

What NVIDIA says (1)

“Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

Practice 2.4 (4 questions) Full Machine Learning With RAPIDS guide

← 2.3 Evaluating and comparing models · 2.5 Cross-validation →