2.1 Training on GPUs with cuML and XGBoost
The scikit-learn style API, cuml.accel and CPU fallback, and XGBoost.
Key points
An estimator is a model object with fit() and predict() methods. cuML keeps the scikit-learn style, so code changes are small.
What NVIDIA says (1)
“cuML estimators look and feel just like scikit-learn estimators .”
cuml.accel intercepts supported scikit-learn calls and sends them to GPU code. It must be turned on before scikit-learn is imported.
What NVIDIA says (2)
“When running a script, use the cuml.accel command-line interface: python -m cuml.accel script.py In Jupyter or IPython, load the extension before other imports: % load_ext cuml.accel”
“Enable cuml.accel before importing scikit-learn, UMAP, or HDBSCAN.”
A fallback runs an unsupported call on the CPU so the code still works. Each switch moves data and time, so many switches eat into the gain.
What NVIDIA says (1)
“CPU fallback preserves compatibility, but frequent transitions between CPU and GPU execution may reduce the overall speedup.”
Gradient-boosted decision trees build many small trees, each fixing errors of the ones before. XGBoost can train on GPUs and scale with Dask or Spark. XGBoost means eXtreme Gradient Boosting.
What NVIDIA says (2)
“XGBoost , which stands for Extreme Gradient Boosting, is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”
“In addition, XGBoost is integrated with distributed processing frameworks like Apache Spark and Dask.”
Key terms
- CPU fallback: Running an operation on the CPU when the GPU path does not support it, at the cost of extra data copies.
- cuML: A GPU machine learning library whose estimators work like scikit-learn estimators.
- cuml.accel: A cuML mode that runs existing scikit-learn code on the GPU without code changes.
- scikit-learn: A popular CPU machine learning library for Python with a standard fit and predict API.
- XGBoost: A scalable library for gradient-boosted decision trees that can train on GPUs.
Sample question
What is cuML, and why is moving scikit-learn code to it usually easy?
Show the answer
Answer: A GPU machine learning library whose estimators look and feel like scikit-learn estimators
An estimator is a model object with fit() and predict() methods. cuML keeps the scikit-learn style, so code changes are small.
What NVIDIA says (1)
“cuML estimators look and feel just like scikit-learn estimators .”
Practice 2.1 (4 questions) Full Machine Learning With RAPIDS guide
← 1.7 Efficient storage with Parquet and modern frameworks · 2.2 Regression, classification and clustering →