3.1 Designing an end-to-end pipeline

NCA-ADS · Data Science Pipelines and Workflow Automation (13% of the exam) · Official objective: “End-to-end data science pipeline design”

Pipeline stages, keeping data on the GPU, and separate reusable flows.

Key points

  1. A pipeline is a chain of steps that turns raw data into results. RAPIDS keeps the data on the GPU across these steps. ETL means extract, transform, load.

    What NVIDIA says (1)

    “RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

    — RAPIDS Accelerates Data Science End-to-End

  2. Serialization means converting data to bytes to move it between tools. Avoiding it between steps saves time and copies. ETL means extract, transform, load. ML means machine learning. API means application programming interface.

    What NVIDIA says (1)

    “This API integrates with a variety of machine learning algorithms without paying typical serialization costs, enabling acceleration for end-to-end pipelines.”

    — RAPIDS Accelerates Data Science End-to-End

  3. Orchestration means running pipeline steps in the right order and handling failures. Modular stages are easier to test, rerun and scale. MLOps means machine learning operations.

    What NVIDIA says (2)

    “Each stage can also run independently for ad-hoc experiments.”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

    “This separation is important for scaling: tasks can be distributed across workers and work pools , while flows define the orchestration logic.”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Key terms

Sample question

Which order best describes an end-to-end data science pipeline that RAPIDS aims to accelerate?

Show the answer

Answer: Data loading, ETL, model training, then inference

A pipeline is a chain of steps that turns raw data into results. RAPIDS keeps the data on the GPU across these steps. ETL means extract, transform, load.

What NVIDIA says (1)

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

Practice 3.1 (3 questions) Full Data Science Pipelines and Workflow Automation guide

← 2.6 Metrics and the confusion matrix · 3.2 Feature selection and transformation →