1.3 GPU-accelerated ETL with RAPIDS, Dask and Spark

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “GPU-accelerated ETL workflows with RAPIDS, Dask, or Spark”

What ETL is, Dask cuDF for data larger than one GPU, and the RAPIDS Accelerator for Apache Spark.

Key points

  1. ETL is the data preparation stage before analysis or training. RAPIDS accelerates it on GPUs.

    What NVIDIA says (2)

    “These tasks are often grouped under the term Extract, Transform, Load (ETL).”

    — RAPIDS Accelerates Data Science End-to-End

    “RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

    — RAPIDS Accelerates Data Science End-to-End

  2. Dask is a Python library for parallel computing. Dask cuDF registers cuDF as the DataFrame backend, so each partition is a cuDF DataFrame. To use several GPUs you also need a dask.distributed cluster; Dask-CUDA makes that easy. CSV means comma-separated values.

    What NVIDIA says (3)

    “When installed, Dask cuDF is automatically registered as the "cudf" dataframe backend for Dask DataFrame .”

    — Dask cuDF documentation

    “You must also deploy a dask.distributed cluster to leverage multiple GPUs.”

    — Dask cuDF documentation

    “Therefore to solve this with the libraries we also need to use Dask to partition the dataset and stream it through GPU memory, then cuDF can process each partition in a performant way.”

    — RAPIDS Deployment: One Billion Row Challenge on a single node

  3. Apache Spark is a distributed engine for large-scale data processing. The RAPIDS Accelerator plugs into Spark and runs supported operations on GPUs. ETL means extract, transform, load. SQL means Structured Query Language.

    What NVIDIA says (2)

    “Run your existing Apache Spark applications with no code change.”

    — cuDF for Apache Spark: Overview

    “The cuDF for Apache Spark combines the power of the RAPIDS cuDF library and the scale of the Spark distributed computing framework.”

    — cuDF for Apache Spark: Overview

  4. Feature engineering means building model inputs from raw data. AT&T's case shows GPUs can pay off in early pipeline stages too, not only in training. ETL means extract, transform, load.

    What NVIDIA says (2)

    “The optimization came from using the RAPIDS accelerator for Apache Spark, an open-source library that enables GPU-accelerated ETL and feature engineering.”

    — Scaling Data Pipelines: AT&T Optimizes Speed, Cost, and Efficiency with GPUs

    “SPOILER ALERT: We were pleasantly surprised that, at least for the examples examined, the use of GPUs for each pipeline stage proved to be faster, cheaper, and simpler!”

    — Scaling Data Pipelines: AT&T Optimizes Speed, Cost, and Efficiency with GPUs

  5. Spilling means moving data from GPU memory to host (CPU) memory to free space. Dask-CUDA has several spilling options; cuDF-level spilling often works best for tables. ETL means extract, transform, load.

    What NVIDIA says (1)

    “We’ve found that for tabular-based workloads, using enable-cudf-spill is often faster and more stable compared with the other Dask-CUDA options, including --device-memory-limit .”

    — Best Practices for Multi-GPU Data Analysis Using RAPIDS with Dask

Key terms

Sample question

What is ETL?

Show the answer

Answer: Extract, transform, load: the steps that pull raw data, clean and reshape it, and store it for analysis

ETL is the data preparation stage before analysis or training. RAPIDS accelerates it on GPUs.

What NVIDIA says (2)

“These tasks are often grouped under the term Extract, Transform, Load (ETL).”

— RAPIDS Accelerates Data Science End-to-End

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

Practice 1.3 (5 questions) Full Data Manipulation and Preparation guide

← 1.2 Cleaning data and handling quality and governance · 1.4 Feature engineering for numbers and categories →