Machine Learning With RAPIDS

16% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

2.1 Training on GPUs with cuML and XGBoost

Official objective: “GPU-accelerated model training with NVIDIA cuML and XGBoost”

The scikit-learn style API, cuml.accel and CPU fallback, and XGBoost.

Key points

  1. An estimator is a model object with fit() and predict() methods. cuML keeps the scikit-learn style, so code changes are small.

    What NVIDIA says (1)

    “cuML estimators look and feel just like scikit-learn estimators .”

    — cuML: Introduction

  2. cuml.accel intercepts supported scikit-learn calls and sends them to GPU code. It must be turned on before scikit-learn is imported.

    What NVIDIA says (2)

    “When running a script, use the cuml.accel command-line interface: python -m cuml.accel script.py In Jupyter or IPython, load the extension before other imports: % load_ext cuml.accel”

    — cuML: cuml.accel usage

    “Enable cuml.accel before importing scikit-learn, UMAP, or HDBSCAN.”

    — cuML: cuml.accel usage

  3. A fallback runs an unsupported call on the CPU so the code still works. Each switch moves data and time, so many switches eat into the gain.

    What NVIDIA says (1)

    “CPU fallback preserves compatibility, but frequent transitions between CPU and GPU execution may reduce the overall speedup.”

    — cuML: cuml.accel usage

  4. Gradient-boosted decision trees build many small trees, each fixing errors of the ones before. XGBoost can train on GPUs and scale with Dask or Spark. XGBoost means eXtreme Gradient Boosting.

    What NVIDIA says (2)

    “XGBoost , which stands for Extreme Gradient Boosting, is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”

    — What Is XGBoost and Why Does It Matter?

    “In addition, XGBoost is integrated with distributed processing frameworks like Apache Spark and Dask.”

    — What Is XGBoost and Why Does It Matter?

Key terms: CPU fallback cuML cuml.accel scikit-learn XGBoost

Practice 2.1 (4 questions) Objective page

2.2 Regression, classification and clustering

Official objective: “Regression, classification, and clustering techniques”

Picking a model family: linear and logistic regression, k-means, DBSCAN and random forests.

Key points

  1. Regression predicts a continuous number. Classification predicts a category. Clustering groups unlabeled data. PCA means principal component analysis.

    What NVIDIA says (1)

    “Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

  2. Logistic regression is a classification model that outputs a probability for each class. UMAP means Uniform Manifold Approximation and Projection.

    What NVIDIA says (1)

    “Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

  3. Unsupervised learning finds patterns in data without labels. Clustering is its most common task. DBSCAN means Density-Based Spatial Clustering of Applications with Noise.

    What NVIDIA says (2)

    “Unsupervised learning algorithms attempt to ‘learn’ patterns in unlabeled data sets, discovering similarities, or regularities.”

    — NVIDIA Glossary: K-Means

    “Common unsupervised tasks include clustering and association.”

    — NVIDIA Glossary: K-Means

  4. k-means assigns every point to the nearest of k centers. DBSCAN instead grows clusters from dense regions and can mark sparse points as noise. DBSCAN means Density-Based Spatial Clustering of Applications with Noise.

    What NVIDIA says (2)

    “DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”

    — cuML API: DBSCAN

    “This also allows DBSCAN to be robust to noise.”

    — cuML API: DBSCAN

  5. Overfitting means a model learns noise in the training data and does worse on new data. Random forests average many varied trees, which smooths out the noise.

    What NVIDIA says (2)

    “Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”

    — NVIDIA Glossary: Random Forest

    “Whereas decision trees are based upon a fixed set of features, and often overfit, randomness is critical to the success of the forest.”

    — NVIDIA Glossary: Random Forest

Key terms: Supervised learning Unsupervised learning Regression Classification Clustering DBSCAN Random forest

Practice 2.2 (5 questions) Objective page

2.3 Evaluating and comparing models

Official objective: “Model evaluation, comparison, and generalization assessment”

Held-out test sets, R-squared, and judging generalization.

Key points

  1. Generalization is how well a model does on data it has not seen. A held-out test set measures it; training accuracy alone can hide overfitting.

    What NVIDIA says (2)

    “test_size float or int, default=None If float, should be between 0.0 and 1.0 and represent the proportion of the dataset to include”

    — cuML API: train_test_split

    “Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  2. R², the coefficient of determination, is the share of the target's variance that a model explains. It is a relative metric for models trained on the same data.

    What NVIDIA says (2)

    “R-squared (R²), also known as the coefficient of determination, represents the proportion of variance explained by a model.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “R² is a relative metric; that is, it can be used to compare with other models trained on the same dataset.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  3. A model that always predicts the mean gets R² = 0. A negative score means the model is worse than that simple baseline.

    What NVIDIA says (1)

    “Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”

    — cuML API: r2_score

  4. Pick the metric that fits the task. Accuracy is the ratio of correct predictions to all predictions; R² fits continuous targets.

    What NVIDIA says (2)

    “Accuracy score is the ratio of correct predictions to the total number of predictions.”

    — cuML: Training and evaluating machine learning models

    “The cell below uses the Linear Regression model and evaluates its performance using cuML’s R² score metric.”

    — cuML: Training and evaluating machine learning models

Key terms: Test set R-squared

Try it: Regression metrics lab

Practice 2.3 (4 questions) Objective page

2.4 Hyperparameter tuning

Official objective: “Hyperparameter tuning and optimization”

Grid search, random search and Optuna with cuML.

Key points

  1. A hyperparameter is a setting chosen before training, unlike weights learned during training. HPO searches for good settings.

    What NVIDIA says (1)

    “Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  2. Grid search tries every combination of the values you list. That is 3 x 3 = 9 here, and the grid grows fast as you add parameters.

    What NVIDIA says (2)

    “The grid search will take place over |n_estimators| x |max_depth| which is 3 x 3 = 9.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

    “As you have probably guessed, the grid size grows rapidly as the number of parameters and their search space increases.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  3. Random search draws a fixed number of combinations at random. In NVIDIA's notebook it matched grid search with just 25 combinations.

    What NVIDIA says (2)

    “Random Search replaces the exhaustive nature of the search from before with a random selection of parameters over the specified space.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

    “We notice that performing grid search and random search yields similar performance improvements even though random search used just 25 combination of parameters.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  4. Optuna is a lightweight framework for automatic HPO. You give it an objective function that trains and scores a model; it chooses the next settings. HPO means hyperparameter optimization.

    What NVIDIA says (2)

    “Optuna is a lightweight framework for automatic hyperparameter optimization.”

    — RAPIDS Deployment: Optuna HPO with RAPIDS

    “By simply wrapping the objective function with Optuna, we can perform a parallel-distributed HPO search over a search space as we’ll see in this notebook.”

    — RAPIDS Deployment: Optuna HPO with RAPIDS

Key terms: Optuna Hyperparameter Hyperparameter optimization Grid search Random search

Practice 2.4 (4 questions) Objective page

2.5 Cross-validation

Official objective: “Cross-validation methods”

k-fold cross-validation, why it helps tuning, and when shuffling matters.

Key points

  1. Cross-validation gives a steadier estimate of performance than one split, because every row is used for validation once.

    What NVIDIA says (1)

    “Each fold is then used once as a validation set while the k - 1 remaining folds form the training set.”

    — cuML API: KFold

  2. During HPO you compare many candidates. Cross-validation reduces the chance of picking settings that only looked good on one lucky split. HPO means hyperparameter optimization.

    What NVIDIA says (2)

    “Cross-validation is often used to more accurately estimate the performance of the models in the search process.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

    “Cross-validation is the method of splitting the training set into complementary subsets and performing training on one of the subsets, then predicting the models performance on the other.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

  3. Without shuffling, each fold is a consecutive block of rows. If the file is sorted, folds may not represent the whole data. For time series, keep order on purpose.

    What NVIDIA says (1)

    “Split dataset into k consecutive folds (without shuffling by default).”

    — cuML API: KFold

Key terms: Cross-validation k-fold cross-validation

Try it: Split lab

Practice 2.5 (3 questions) Objective page

2.6 Metrics and the confusion matrix

Official objective: “Performance metrics and confusion matrix interpretation”

Precision, recall, the confusion matrix, and robust regression metrics.

Key points

  1. Precision = true positives / (true positives + false positives) = 80/100. Recall = true positives / (true positives + false negatives) = 80/200.

    What NVIDIA says (2)

    “The precision is the ratio tp / (tp + fp) where tp is the number of true positives and fp the number of false positives.”

    — cuML API: precision_recall_curve

    “The recall is the ratio tp / (tp + fn) where tp is the number of true positives and fn the number of false negatives.”

    — cuML API: precision_recall_curve

  2. A confusion matrix counts how often each true class was predicted as each class. Off-diagonal cells are mistakes.

    What NVIDIA says (1)

    “Normalizes confusion matrix over the true (rows), predicted (columns) conditions or al”

    — cuML API: confusion_matrix

  3. Recall measures how many positives you catch. When misses are expensive, high recall matters most, even at some cost in precision.

    What NVIDIA says (1)

    “The recall is intuitively the ability of the classifier to find all the positive samples.”

    — cuML API: precision_recall_curve

  4. MSE squares each error, so a few big errors dominate it. MAE averages absolute errors, so it is more robust to outliers.

    What NVIDIA says (2)

    “Some of those might be outliers, so MSE is not robust to their presence.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “Other than the scale, RMSE has the same properties as MSE.”

    — A Comprehensive Overview of Regression Evaluation Metrics

Key terms: Accuracy Precision Recall Confusion matrix MSE and RMSE Mean absolute error

Try it: Confusion matrix lab Regression metrics lab

Practice 2.6 (4 questions) Objective page