3.6 Reproducible pipelines with RAPIDS and Dask

NCA-ADS · Data Science Pipelines and Workflow Automation (13% of the exam) · Official objective: “Building reproducible pipelines with RAPIDS and Dask”

LocalCUDACluster, environment files and floating-point determinism.

Key points

  1. A Dask worker is a process that runs tasks. LocalCUDACluster pins one worker to each GPU on the machine.

    What NVIDIA says (2)

    “For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”

    — cuML: Multi-GPU with Dask

    “This automatically creates one worker per available GPU.”

    — cuML: Multi-GPU with Dask

  2. Reproducible means someone else can rebuild the same environment and get the same results. An environment file pins the packages. MLOps means machine learning operations.

    What NVIDIA says (1)

    “All dependencies (CUDA-X libraries, Prefect, MLflow, Triton client) are specified in environment.yml : $ conda env create -f environment.yml”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  3. Floating-point addition can give slightly different answers in a different order. Reproducible pipelines compare results up to a chosen precision.

    What NVIDIA says (2)

    “This impacts the determinism of floating-point operations because floating-point arithmetic is non-associative, that is, (a + b) + c is not necessarily equal to a + (b + c) .”

    — cuDF: Comparison of cuDF and pandas

    “If you need to compare floating point results, you should typically do so using the functions provided in the cudf.testing module, which allow you to compare values up to a desired precision.”

    — cuDF: Comparison of cuDF and pandas

Key terms

Sample question

How do you set up a single-node, multi-GPU Dask cluster for cuML?

Show the answer

Answer: Use LocalCUDACluster from dask-cuda, which creates one worker per available GPU

A Dask worker is a process that runs tasks. LocalCUDACluster pins one worker to each GPU on the machine.

What NVIDIA says (2)

“For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”

— cuML: Multi-GPU with Dask

“This automatically creates one worker per available GPU.”

— cuML: Multi-GPU with Dask

Practice 3.6 (3 questions) Full Data Science Pipelines and Workflow Automation guide

← 3.5 Automating and scaling workflows · 4.1 Exploratory data analysis and descriptive statistics →