3.6 Reproducible pipelines with RAPIDS and Dask
LocalCUDACluster, environment files and floating-point determinism.
Key points
A Dask worker is a process that runs tasks. LocalCUDACluster pins one worker to each GPU on the machine.
What NVIDIA says (2)
“For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”
“This automatically creates one worker per available GPU.”
Reproducible means someone else can rebuild the same environment and get the same results. An environment file pins the packages. MLOps means machine learning operations.
What NVIDIA says (1)
“All dependencies (CUDA-X libraries, Prefect, MLflow, Triton client) are specified in environment.yml : $ conda env create -f environment.yml”
Floating-point addition can give slightly different answers in a different order. Reproducible pipelines compare results up to a chosen precision.
What NVIDIA says (2)
“This impacts the determinism of floating-point operations because floating-point arithmetic is non-associative, that is, (a + b) + c is not necessarily equal to a + (b + c) .”
“If you need to compare floating point results, you should typically do so using the functions provided in the cudf.testing module, which allow you to compare values up to a desired precision.”
Key terms
- LocalCUDACluster: A Dask-CUDA cluster on one machine with one worker per GPU.
- Floating-point determinism: Getting bit-identical results on every run; parallel float sums may differ slightly, so compare with a tolerance.
- Environment file: A text file, such as env.yaml or requirements.txt, that lists the packages a project needs.
Sample question
How do you set up a single-node, multi-GPU Dask cluster for cuML?
Show the answer
Answer: Use LocalCUDACluster from dask-cuda, which creates one worker per available GPU
A Dask worker is a process that runs tasks. LocalCUDACluster pins one worker to each GPU on the machine.
What NVIDIA says (2)
“For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”
“This automatically creates one worker per available GPU.”
Practice 3.6 (3 questions) Full Data Science Pipelines and Workflow Automation guide
← 3.5 Automating and scaling workflows · 4.1 Exploratory data analysis and descriptive statistics →