Data Science Pipelines and Workflow Automation
13% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management
3.1 Designing an end-to-end pipeline
Pipeline stages, keeping data on the GPU, and separate reusable flows.
Key points
A pipeline is a chain of steps that turns raw data into results. RAPIDS keeps the data on the GPU across these steps. ETL means extract, transform, load.
What NVIDIA says (1)
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
Serialization means converting data to bytes to move it between tools. Avoiding it between steps saves time and copies. ETL means extract, transform, load. ML means machine learning. API means application programming interface.
What NVIDIA says (1)
“This API integrates with a variety of machine learning algorithms without paying typical serialization costs, enabling acceleration for end-to-end pipelines.”
Orchestration means running pipeline steps in the right order and handling failures. Modular stages are easier to test, rerun and scale. MLOps means machine learning operations.
What NVIDIA says (2)
“Each stage can also run independently for ad-hoc experiments.”
“This separation is important for scaling: tasks can be distributed across workers and work pools , while flows define the orchestration logic.”
Key terms: RAPIDS Data science pipeline
3.2 Feature selection and transformation
Column transformers, scaling and encoding, feature importance and SHAP.
Key points
A transformer changes features, for example by scaling or encoding. ColumnTransformer runs a chosen transformer on each column group and concatenates the outputs. PCA means principal component analysis.
What NVIDIA says (1)
“This estimator allows different columns or column subsets of the input to be transformed separately and the features generated by each transformer will be concatenated to form a single feature space.”
Feature transformation reshapes inputs so a model can use them. Normalizing fixes scale differences; extracting the hour exposes a daily pattern.
What NVIDIA says (2)
“Next, numeric features must be normalized to prevent the model from being biased by the variable scales.”
“Then transform the ‘date’ column into an ‘hour’ feature, as weather patterns often correlate with the time of day.”
Feature selection keeps the inputs that help and drops the rest. Impurity-based importances from a forest are a common first guide.
What NVIDIA says (1)
“feature_importances_ ndarray of shape (n_features,) The impurity-based feature importances.”
SHAP values assign each feature a share of a prediction. They help select features and explain models. SHAP means SHapley Additive exPlanations.
What NVIDIA says (1)
“cuML’s SHAP based explainers accelerate the algorithmic part of SHAP.”
Key terms: Feature Feature engineering Standard scaling Feature importance SHAP
Try it: Feature lab
3.3 Fixing underfitting and overfitting
Bagging versus boosting, regularization and simpler models.
Key points
Overfitting means the model learned noise in the training data. Regularization penalizes complexity so the model generalizes better.
What NVIDIA says (2)
“Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”
“Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”
Variance error comes from a model that changes too much with the data. Bias error comes from a model that is too simple. GBDT means gradient-boosted decision trees.
What NVIDIA says (1)
“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”
Regularization adds a penalty for large weights. Ridge uses an L2 penalty; Lasso uses an L1 penalty that can set weights to zero.
What NVIDIA says (2)
“Ridge extends LinearRegression by providing L2 regularization on the coefficients when predicting response y with a linear combination of the predictors in X.”
“Larger values specify stronger regularization.”
Key terms: Random forest Bagging Boosting Regularization Overfitting Underfitting
3.4 Augmenting and integrating datasets
Synthetic data for rare cases, derived features and safe joins.
Key points
Data augmentation adds new training examples based on existing data or simulations.
What NVIDIA says (2)
“Synthetically generated data can augment existing data for larger, more representative datasets.”
“If data used to train self-driving cars underrepresents uncommon scenes such as extreme weather conditions or traffic accidents, synthetic data can help augment the diversity of these datasets to better”
Integration means combining data or results from several sources into one dataset. Adding a computed column is a simple form of it. EDA means exploratory data analysis.
What NVIDIA says (1)
“Augmenting the data using RAPIDS cuSpatial to quickly calculate distances also shows that most trips are relatively short.”
Integrating data with a merge can add rows when keys repeat and lose rows when keys are missing. In cuDF, a left join keeps every row of the left table.
What NVIDIA says (1)
“df_merged = df_a . merge ( df_b , on = [ 'key' ], how = 'left' )”
Key terms: Join (merge) Synthetic data Data augmentation
3.5 Automating and scaling workflows
Orchestration with Prefect, multi-GPU Dask and the data flywheel.
Key points
Automation means the pipeline runs on a schedule or trigger without manual steps. MLOps means machine learning operations.
What NVIDIA says (1)
“Prefect orchestrates the pipeline stages, tracks runs, and enables scheduling”
Scalability means handling more data or work by adding resources. cuML's Dask estimators split the work across GPUs.
What NVIDIA says (1)
“While cuML’s single-GPU implementations are highly optimized, distributed computing with Dask enables you to: Scale beyond single GPU memory : Process datasets larger than what fits on a single GPU Accelerate training : Distribute computation across multiple GPUs for faster model training”
A data flywheel turns usage into better models over time. Each cycle collects data, refines the model, evaluates and redeploys.
What NVIDIA says (2)
“AI data flywheels work by creating a loop where AI models continuously improve by learning from the latest institutional knowledge and user feedback.”
“As the system interacts with the environment, it collects feedback and new data, which are then used to refine and enhance the backbone models powering the AI workflows.”
Key terms: LocalCUDACluster Prefect Data flywheel
3.6 Reproducible pipelines with RAPIDS and Dask
LocalCUDACluster, environment files and floating-point determinism.
Key points
A Dask worker is a process that runs tasks. LocalCUDACluster pins one worker to each GPU on the machine.
What NVIDIA says (2)
“For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”
“This automatically creates one worker per available GPU.”
Reproducible means someone else can rebuild the same environment and get the same results. An environment file pins the packages. MLOps means machine learning operations.
What NVIDIA says (1)
“All dependencies (CUDA-X libraries, Prefect, MLflow, Triton client) are specified in environment.yml : $ conda env create -f environment.yml”
Floating-point addition can give slightly different answers in a different order. Reproducible pipelines compare results up to a chosen precision.
What NVIDIA says (2)
“This impacts the determinism of floating-point operations because floating-point arithmetic is non-associative, that is, (a + b) + c is not necessarily equal to a + (b + c) .”
“If you need to compare floating point results, you should typically do so using the functions provided in the cudf.testing module, which allow you to compare values up to a desired precision.”
Key terms: LocalCUDACluster Floating-point determinism Environment file