3.5 Automating and scaling workflows

NCA-ADS · Data Science Pipelines and Workflow Automation (13% of the exam) · Official objective: “Automation and scalability of data science workflows”

Orchestration with Prefect, multi-GPU Dask and the data flywheel.

Key points

  1. Automation means the pipeline runs on a schedule or trigger without manual steps. MLOps means machine learning operations.

    What NVIDIA says (1)

    “Prefect orchestrates the pipeline stages, tracks runs, and enables scheduling”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  2. Scalability means handling more data or work by adding resources. cuML's Dask estimators split the work across GPUs.

    What NVIDIA says (1)

    “While cuML’s single-GPU implementations are highly optimized, distributed computing with Dask enables you to: Scale beyond single GPU memory : Process datasets larger than what fits on a single GPU Accelerate training : Distribute computation across multiple GPUs for faster model training”

    — cuML: Multi-GPU with Dask

  3. A data flywheel turns usage into better models over time. Each cycle collects data, refines the model, evaluates and redeploys.

    What NVIDIA says (2)

    “AI data flywheels work by creating a loop where AI models continuously improve by learning from the latest institutional knowledge and user feedback.”

    — Data flywheel: What it is and how it works

    “As the system interacts with the environment, it collects feedback and new data, which are then used to refine and enhance the backbone models powering the AI workflows.”

    — Data flywheel: What it is and how it works

Key terms

Sample question

In the fraud MLOps example, which tool orchestrates the pipeline, tracks runs and handles scheduling?

Show the answer

Answer: Prefect

Automation means the pipeline runs on a schedule or trigger without manual steps. MLOps means machine learning operations.

What NVIDIA says (1)

“Prefect orchestrates the pipeline stages, tracks runs, and enables scheduling”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Practice 3.5 (3 questions) Full Data Science Pipelines and Workflow Automation guide

← 3.4 Augmenting and integrating datasets · 3.6 Reproducible pipelines with RAPIDS and Dask →