Data Science Pipelines and Workflow Automation

13% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

3.1 Designing an end-to-end pipeline

Official objective: “End-to-end data science pipeline design”

Pipeline stages, keeping data on the GPU, and separate reusable flows.

Key points

  1. A pipeline is a chain of steps that turns raw data into results. RAPIDS keeps the data on the GPU across these steps. ETL means extract, transform, load.

    What NVIDIA says (1)

    “RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

    — RAPIDS Accelerates Data Science End-to-End

  2. Serialization means converting data to bytes to move it between tools. Avoiding it between steps saves time and copies. ETL means extract, transform, load. ML means machine learning. API means application programming interface.

    What NVIDIA says (1)

    “This API integrates with a variety of machine learning algorithms without paying typical serialization costs, enabling acceleration for end-to-end pipelines.”

    — RAPIDS Accelerates Data Science End-to-End

  3. Orchestration means running pipeline steps in the right order and handling failures. Modular stages are easier to test, rerun and scale. MLOps means machine learning operations.

    What NVIDIA says (2)

    “Each stage can also run independently for ad-hoc experiments.”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

    “This separation is important for scaling: tasks can be distributed across workers and work pools , while flows define the orchestration logic.”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Key terms: RAPIDS Data science pipeline

Practice 3.1 (3 questions) Objective page

3.2 Feature selection and transformation

Official objective: “Feature engineering, selection, and transformation for model improvement”

Column transformers, scaling and encoding, feature importance and SHAP.

Key points

  1. A transformer changes features, for example by scaling or encoding. ColumnTransformer runs a chosen transformer on each column group and concatenates the outputs. PCA means principal component analysis.

    What NVIDIA says (1)

    “This estimator allows different columns or column subsets of the input to be transformed separately and the features generated by each transformer will be concatenated to form a single feature space.”

    — cuML API: ColumnTransformer

  2. Feature transformation reshapes inputs so a model can use them. Normalizing fixes scale differences; extracting the hour exposes a daily pattern.

    What NVIDIA says (2)

    “Next, numeric features must be normalized to prevent the model from being biased by the variable scales.”

    — Accelerated Data Analytics: Machine Learning with GPU-Accelerated pandas and scikit-learn

    “Then transform the ‘date’ column into an ‘hour’ feature, as weather patterns often correlate with the time of day.”

    — Accelerated Data Analytics: Machine Learning with GPU-Accelerated pandas and scikit-learn

  3. Feature selection keeps the inputs that help and drops the rest. Impurity-based importances from a forest are a common first guide.

    What NVIDIA says (1)

    “feature_importances_ ndarray of shape (n_features,) The impurity-based feature importances.”

    — cuML API: RandomForestClassifier

  4. SHAP values assign each feature a share of a prediction. They help select features and explain models. SHAP means SHapley Additive exPlanations.

    What NVIDIA says (1)

    “cuML’s SHAP based explainers accelerate the algorithmic part of SHAP.”

    — cuML API: KernelExplainer

Key terms: Feature Feature engineering Standard scaling Feature importance SHAP

Try it: Feature lab

Practice 3.2 (4 questions) Objective page

3.3 Fixing underfitting and overfitting

Official objective: “Mitigating underfitting and overfitting through model and feature adjustments”

Bagging versus boosting, regularization and simpler models.

Key points

  1. Overfitting means the model learned noise in the training data. Regularization penalizes complexity so the model generalizes better.

    What NVIDIA says (2)

    “Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”

    — Gradient Boosting, Decision Trees and XGBoost with CUDA

    “Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”

    — Gradient Boosting, Decision Trees and XGBoost with CUDA

  2. Variance error comes from a model that changes too much with the data. Bias error comes from a model that is too simple. GBDT means gradient-boosted decision trees.

    What NVIDIA says (1)

    “Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”

    — NVIDIA Glossary: Random Forest

  3. Regularization adds a penalty for large weights. Ridge uses an L2 penalty; Lasso uses an L1 penalty that can set weights to zero.

    What NVIDIA says (2)

    “Ridge extends LinearRegression by providing L2 regularization on the coefficients when predicting response y with a linear combination of the predictors in X.”

    — cuML API: Ridge

    “Larger values specify stronger regularization.”

    — cuML API: Ridge

Key terms: Random forest Bagging Boosting Regularization Overfitting Underfitting

Practice 3.3 (3 questions) Objective page

3.4 Augmenting and integrating datasets

Official objective: “Dataset augmentation and integration for enhanced training data”

Synthetic data for rare cases, derived features and safe joins.

Key points

  1. Data augmentation adds new training examples based on existing data or simulations.

    What NVIDIA says (2)

    “Synthetically generated data can augment existing data for larger, more representative datasets.”

    — NVIDIA Glossary: Synthetic Data Generation

    “If data used to train self-driving cars underrepresents uncommon scenes such as extreme weather conditions or traffic accidents, synthetic data can help augment the diversity of these datasets to better”

    — What Is Trustworthy AI?

  2. Integration means combining data or results from several sources into one dataset. Adding a computed column is a simple form of it. EDA means exploratory data analysis.

    What NVIDIA says (1)

    “Augmenting the data using RAPIDS cuSpatial to quickly calculate distances also shows that most trips are relatively short.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. Integrating data with a merge can add rows when keys repeat and lose rows when keys are missing. In cuDF, a left join keeps every row of the left table.

    What NVIDIA says (1)

    “df_merged = df_a . merge ( df_b , on = [ 'key' ], how = 'left' )”

    — cuDF API: DataFrame.merge

Key terms: Join (merge) Synthetic data Data augmentation

Practice 3.4 (3 questions) Objective page

3.5 Automating and scaling workflows

Official objective: “Automation and scalability of data science workflows”

Orchestration with Prefect, multi-GPU Dask and the data flywheel.

Key points

  1. Automation means the pipeline runs on a schedule or trigger without manual steps. MLOps means machine learning operations.

    What NVIDIA says (1)

    “Prefect orchestrates the pipeline stages, tracks runs, and enables scheduling”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  2. Scalability means handling more data or work by adding resources. cuML's Dask estimators split the work across GPUs.

    What NVIDIA says (1)

    “While cuML’s single-GPU implementations are highly optimized, distributed computing with Dask enables you to: Scale beyond single GPU memory : Process datasets larger than what fits on a single GPU Accelerate training : Distribute computation across multiple GPUs for faster model training”

    — cuML: Multi-GPU with Dask

  3. A data flywheel turns usage into better models over time. Each cycle collects data, refines the model, evaluates and redeploys.

    What NVIDIA says (2)

    “AI data flywheels work by creating a loop where AI models continuously improve by learning from the latest institutional knowledge and user feedback.”

    — Data flywheel: What it is and how it works

    “As the system interacts with the environment, it collects feedback and new data, which are then used to refine and enhance the backbone models powering the AI workflows.”

    — Data flywheel: What it is and how it works

Key terms: LocalCUDACluster Prefect Data flywheel

Practice 3.5 (3 questions) Objective page

3.6 Reproducible pipelines with RAPIDS and Dask

Official objective: “Building reproducible pipelines with RAPIDS and Dask”

LocalCUDACluster, environment files and floating-point determinism.

Key points

  1. A Dask worker is a process that runs tasks. LocalCUDACluster pins one worker to each GPU on the machine.

    What NVIDIA says (2)

    “For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”

    — cuML: Multi-GPU with Dask

    “This automatically creates one worker per available GPU.”

    — cuML: Multi-GPU with Dask

  2. Reproducible means someone else can rebuild the same environment and get the same results. An environment file pins the packages. MLOps means machine learning operations.

    What NVIDIA says (1)

    “All dependencies (CUDA-X libraries, Prefect, MLflow, Triton client) are specified in environment.yml : $ conda env create -f environment.yml”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  3. Floating-point addition can give slightly different answers in a different order. Reproducible pipelines compare results up to a chosen precision.

    What NVIDIA says (2)

    “This impacts the determinism of floating-point operations because floating-point arithmetic is non-associative, that is, (a + b) + c is not necessarily equal to a + (b + c) .”

    — cuDF: Comparison of cuDF and pandas

    “If you need to compare floating point results, you should typically do so using the functions provided in the cudf.testing module, which allow you to compare values up to a desired precision.”

    — cuDF: Comparison of cuDF and pandas

Key terms: LocalCUDACluster Floating-point determinism Environment file

Practice 3.6 (3 questions) Objective page