2.5 Cross-validation

NCA-ADS · Machine Learning With RAPIDS (16% of the exam) · Official objective: “Cross-validation methods”

k-fold cross-validation, why it helps tuning, and when shuffling matters.

Key points

  1. Cross-validation gives a steadier estimate of performance than one split, because every row is used for validation once.

    What NVIDIA says (1)

    “Each fold is then used once as a validation set while the k - 1 remaining folds form the training set.”

    — cuML API: KFold

  2. During HPO you compare many candidates. Cross-validation reduces the chance of picking settings that only looked good on one lucky split. HPO means hyperparameter optimization.

    What NVIDIA says (2)

    “Cross-validation is often used to more accurately estimate the performance of the models in the search process.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

    “Cross-validation is the method of splitting the training set into complementary subsets and performing training on one of the subsets, then predicting the models performance on the other.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  3. Without shuffling, each fold is a consecutive block of rows. If the file is sorted, folds may not represent the whole data. For time series, keep order on purpose.

    What NVIDIA says (1)

    “Split dataset into k consecutive folds (without shuffling by default).”

    — cuML API: KFold

Key terms

Try it

Sample question

How does k-fold cross-validation work?

Show the answer

Answer: Split the data into k folds; train on k-1 folds and validate on the remaining one, k times so each fold is used once

Cross-validation gives a steadier estimate of performance than one split, because every row is used for validation once.

What NVIDIA says (1)

“Each fold is then used once as a validation set while the k - 1 remaining folds form the training set.”

— cuML API: KFold

Practice 2.5 (3 questions) Full Machine Learning With RAPIDS guide

← 2.4 Hyperparameter tuning · 2.6 Metrics and the confusion matrix →