2.1 Training on GPUs with cuML and XGBoost

NCA-ADS · Machine Learning With RAPIDS (16% of the exam) · Official objective: “GPU-accelerated model training with NVIDIA cuML and XGBoost”

The scikit-learn style API, cuml.accel and CPU fallback, and XGBoost.

Key points

  1. An estimator is a model object with fit() and predict() methods. cuML keeps the scikit-learn style, so code changes are small.

    What NVIDIA says (1)

    “cuML estimators look and feel just like scikit-learn estimators .”

    — cuML: Introduction

  2. cuml.accel intercepts supported scikit-learn calls and sends them to GPU code. It must be turned on before scikit-learn is imported.

    What NVIDIA says (2)

    “When running a script, use the cuml.accel command-line interface: python -m cuml.accel script.py In Jupyter or IPython, load the extension before other imports: % load_ext cuml.accel”

    — cuML: cuml.accel usage

    “Enable cuml.accel before importing scikit-learn, UMAP, or HDBSCAN.”

    — cuML: cuml.accel usage

  3. A fallback runs an unsupported call on the CPU so the code still works. Each switch moves data and time, so many switches eat into the gain.

    What NVIDIA says (1)

    “CPU fallback preserves compatibility, but frequent transitions between CPU and GPU execution may reduce the overall speedup.”

    — cuML: cuml.accel usage

  4. Gradient-boosted decision trees build many small trees, each fixing errors of the ones before. XGBoost can train on GPUs and scale with Dask or Spark. XGBoost means eXtreme Gradient Boosting.

    What NVIDIA says (2)

    “XGBoost , which stands for Extreme Gradient Boosting, is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”

    — What Is XGBoost and Why Does It Matter?

    “In addition, XGBoost is integrated with distributed processing frameworks like Apache Spark and Dask.”

    — What Is XGBoost and Why Does It Matter?

Key terms

Sample question

What is cuML, and why is moving scikit-learn code to it usually easy?

Show the answer

Answer: A GPU machine learning library whose estimators look and feel like scikit-learn estimators

An estimator is a model object with fit() and predict() methods. cuML keeps the scikit-learn style, so code changes are small.

What NVIDIA says (1)

“cuML estimators look and feel just like scikit-learn estimators .”

— cuML: Introduction

Practice 2.1 (4 questions) Full Machine Learning With RAPIDS guide

← 1.7 Efficient storage with Parquet and modern frameworks · 2.2 Regression, classification and clustering →