5.4 The end-to-end data science workflow

NCA-ADS · Foundations of Accelerated Data Science (12% of the exam) · Official objective: “End-to-end data science workflow (ingest, ETL, clean, transform)”

Ingest, ETL, clean and transform, and keeping every stage on the GPU.

Key points

  1. Ingest means bringing raw data in. ETL, cleaning and transformation prepare it for training. ETL means extract, transform, load.

    What NVIDIA says (1)

    “RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

    — RAPIDS Accelerates Data Science End-to-End

  2. Cleaning fixes types, gaps and errors so data can be analyzed. cuDF offers the same style of API on GPUs. MLOps means machine learning operations. API means application programming interface.

    What NVIDIA says (1)

    “With its support for structured data formats like tables, matrices, and time series, the pandas Python API provides tools to process messy or raw datasets into clean, structured formats ready for analysis.”

    — What Is Pandas and Why Does it Matter?

  3. Each move between tools or devices costs time. Keeping the whole pipeline on the GPU removes most of those moves. ETL means extract, transform, load.

    What NVIDIA says (1)

    “To get the highest performance for your ML pipeline on NVIDIA GPUs, minimize data transfers between the CPUs and GPUs as part of your pipeline.”

    — NVIDIA cuML Brings Zero Code Change Acceleration to scikit-learn

Key terms

Sample question

Which list matches the early stages of a data science workflow before modeling?

Show the answer

Answer: Ingest data, run ETL, clean it, then transform it into features

Ingest means bringing raw data in. ETL, cleaning and transformation prepare it for training. ETL means extract, transform, load.

What NVIDIA says (1)

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

Practice 5.4 (3 questions) Full Foundations of Accelerated Data Science guide

← 5.3 CPU versus GPU workloads and memory transfers · 5.5 Distributed versus GPU-accelerated frameworks →