Data Manipulation and Preparation
23% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management
1.1 Joining and manipulating data with cuDF and pandas
What cuDF is, cudf.pandas zero-code-change acceleration, joins, group-bys, and why row loops are slow on GPUs.
Key points
A DataFrame is a table with named columns. cuDF holds DataFrames in GPU memory and offers the same style of API as pandas, the most common Python table library. API means application programming interface. CUDA means NVIDIA's parallel computing platform.
What NVIDIA says (1)
“cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”
A join (merge) combines two tables on matching key values. pandas keeps or sorts the key order; cuDF does not, because ordering costs time on a parallel GPU. Sort by the key or index after the merge when order matters.
What NVIDIA says (2)
“By contrast, cuDF’s default behavior is to return rows in a non-deterministic order to maximize performance.”
“Note that the dataframe order is not maintained , but may be restored post-merge by sorting by the index.”
cudf.pandas is cuDF's pandas accelerator mode. You keep your pandas code; supported operations run on the GPU, and anything else falls back to normal pandas. API means application programming interface. SQL means Structured Query Language.
What NVIDIA says (2)
“It supports 100% of the Pandas API , using the GPU for supported operations, and automatically falling back to pandas for other operations.”
“Nothing changes, not even your import statements, when going from CPU to GPU.”
cudf.pandas wraps each object in a proxy, a stand-in that can point to a GPU or a CPU copy. It tries the GPU first and falls back to the CPU only for the failing step.
What NVIDIA says (2)
“Attribute lookups and method calls are first attempted on the GPU (copying from CPU if necessary).”
“If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”
GPUs are built to process many values at once, not one at a time. A vectorized function works on a whole column in one call.
What NVIDIA says (2)
“This is because iterating over data that resides on the GPU will yield extremely poor performance, as GPUs are optimized for highly parallel operations rather than sequential operations.”
“If you absolutely must iterate, copy the data from GPU to CPU by using .to_arrow() or .to_pandas() , then convert the result back to GPU using a Series , DataFrame or Index constructor.”
cuDF mirrors common pandas calls such as read_csv, groupby and agg. For code like this, swapping the import is enough. API means application programming interface. CUDA means NVIDIA's parallel computing platform.
What NVIDIA says (1)
“With RAPIDS, all you have to do to use this code to run on a GPU and enjoy the interactive querying of data is to change the import statement.”
Key terms: pandas cuDF cudf.pandas CPU fallback Join (merge) Group-by aggregation
1.2 Cleaning data and handling quality and governance
Nulls in cuDF, dropna and fillna, duplicates, and protecting personal data.
Key points
A missing value is a cell with no data, also called a null. cuDF allows nulls in every data type and shows them as <NA>. NA means not available (missing).
What NVIDIA says (2)
“cudf supports having missing values in all dtypes.”
“To detect missing values, you can use isna() and notna() functions.”
dropna removes rows or columns with nulls. how='any' drops a row with at least one null; how='all' drops only fully empty rows.
What NVIDIA says (1)
“any (default) drops rows (or columns) containing at least one null value. all drops only rows (or columns) containing all null values.”
Imputation means filling missing values with a chosen value. fillna takes a scalar, a Series, or a dict of per-column values.
What NVIDIA says (1)
“A dict can be used to provide different values to fill nulls in different columns.”
Data governance means rules about who may use which data, and how. Compliance means following those rules and the law, such as privacy rules for personal data.
What NVIDIA says (2)
“In addition to complying with privacy and consumer protection laws, trustworthy AI models are tested for safety, security and mitigation of unwanted bias.”
“Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”
Data quality means data is correct, complete and free of junk such as duplicates. drop_duplicates() keeps the first copy by default and drops the rest.
What NVIDIA says (2)
“Training models on these datasets without proper processing can result in higher training time and lower model quality.”
“Determines which duplicates (if any) to keep. - ‘first’ : Drop duplicates except for the first occurrence.”
Key terms: Null (missing value) dropna fillna Deduplication Personal identifiable information Data governance
1.3 GPU-accelerated ETL with RAPIDS, Dask and Spark
What ETL is, Dask cuDF for data larger than one GPU, and the RAPIDS Accelerator for Apache Spark.
Key points
ETL is the data preparation stage before analysis or training. RAPIDS accelerates it on GPUs.
What NVIDIA says (2)
“These tasks are often grouped under the term Extract, Transform, Load (ETL).”
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
Dask is a Python library for parallel computing. Dask cuDF registers cuDF as the DataFrame backend, so each partition is a cuDF DataFrame. To use several GPUs you also need a dask.distributed cluster; Dask-CUDA makes that easy. CSV means comma-separated values.
What NVIDIA says (3)
“When installed, Dask cuDF is automatically registered as the "cudf" dataframe backend for Dask DataFrame .”
“You must also deploy a dask.distributed cluster to leverage multiple GPUs.”
“Therefore to solve this with the libraries we also need to use Dask to partition the dataset and stream it through GPU memory, then cuDF can process each partition in a performant way.”
Apache Spark is a distributed engine for large-scale data processing. The RAPIDS Accelerator plugs into Spark and runs supported operations on GPUs. ETL means extract, transform, load. SQL means Structured Query Language.
What NVIDIA says (2)
“Run your existing Apache Spark applications with no code change.”
“The cuDF for Apache Spark combines the power of the RAPIDS cuDF library and the scale of the Spark distributed computing framework.”
Feature engineering means building model inputs from raw data. AT&T's case shows GPUs can pay off in early pipeline stages too, not only in training. ETL means extract, transform, load.
What NVIDIA says (2)
“The optimization came from using the RAPIDS accelerator for Apache Spark, an open-source library that enables GPU-accelerated ETL and feature engineering.”
“SPOILER ALERT: We were pleasantly surprised that, at least for the examples examined, the use of GPUs for each pipeline stage proved to be faster, cheaper, and simpler!”
Spilling means moving data from GPU memory to host (CPU) memory to free space. Dask-CUDA has several spilling options; cuDF-level spilling often works best for tables. ETL means extract, transform, load.
What NVIDIA says (1)
“We’ve found that for tabular-based workloads, using enable-cudf-spill is often faster and more stable compared with the other Dask-CUDA options, including --device-memory-limit .”
Key terms: Dask Dask cuDF RAPIDS Accelerator for Apache Spark ETL
1.4 Feature engineering for numbers and categories
One-hot and target encoding, standard scaling and binning.
Key points
One-hot encoding turns a categorical column into one true/false column per category. Categorical means the values are labels, not numbers.
What NVIDIA says (1)
“cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”
Target encoding replaces a category with a statistic of the label, usually its mean. cuML's TargetEncoder uses folds to avoid label leakage, which is when the label sneaks into a feature.
What NVIDIA says (2)
“The input data is grouped by the columns Xs and the aggregated mean value of Y of each group is calculated to replace each value of Xs .”
“Several optimizations are applied to prevent label leakage and parallelize the execution.”
Scaling puts numeric features on a similar range so no feature dominates by its units. StandardScaler computes z = (x - mean) / std and reuses those training statistics later.
What NVIDIA says (2)
“Standardize features by removing the mean and scaling to unit variance”
“Mean and standard deviation are then stored to be used on later data using transform() .”
Binning, also called discretization, groups continuous numbers into intervals. Each bin can then be treated as a category.
What NVIDIA says (1)
“Bin continuous data into intervals.”
With all categories present, one column can be computed from the others. Dropping one removes that redundancy but changes how a penalized model treats categories.
What NVIDIA says (1)
“However, dropping one category breaks the symmetry of the original representation and can therefore induce a bias in downstream models, for instance for penalized linear classification or regression models.”
Key terms: Feature Feature engineering One-hot encoding Target encoding Standard scaling Binning
Try it: Feature lab
1.5 Class imbalance and synthetic data
Why accuracy misleads, class weights, stratified splits and synthetic data.
Key points
Class imbalance means one label is much rarer than another. A balanced class weight makes errors on the rare class cost more during training.
What NVIDIA says (2)
“The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”
“If 'balanced' , class weights are computed from the training labels.”
A stratified split keeps each class's share the same in every part. Without it, a small test set may contain few or no rare cases.
What NVIDIA says (1)
“If not None, data is split in a stratified fashion, using this as the class labels.”
Synthetic data is artificial data made by rules, simulations or generative models. It can fill gaps in rare classes and protect privacy.
What NVIDIA says (2)
“Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”
“Data Privacy : Synthetic data helps overcome privacy issues by generating training data that mimic real-world statistics without directly corresponding to individual records.”
Accuracy is the share of correct predictions. On imbalanced data it hides poor results on the rare class; a confusion matrix shows them.
What NVIDIA says (2)
“Accuracy classification score.”
“Compute confusion matrix to evaluate the accuracy of a classification.”
Key terms: Class imbalance Stratified split Synthetic data Accuracy
Try it: Confusion matrix lab
1.6 Dimensionality reduction and sampling
PCA, UMAP, random projection and reproducible sampling.
Key points
Dimensionality reduction means representing data with fewer columns while keeping most of the information. PCA keeps the directions with the most variance. PCA means principal component analysis.
What NVIDIA says (1)
“PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”
UMAP builds a nearest-neighbor graph and embeds the data into fewer dimensions. It keeps local neighborhoods, which makes clusters visible. UMAP means Uniform Manifold Approximation and Projection.
What NVIDIA says (2)
“UMAP is a popular dimension reduction algorithm used in fields like bioinformatics, NLP topic modeling, and ML preprocessing.”
“It works by creating a k-nearest neighbors (k-NN) graph, which is known in literature as an all-neighbors graph, to build a fuzzy topological representation of the data, which is used to embed high-dimensional data into lower dimensions.”
Sampling picks a subset of rows to work with. A fixed random_state gives the same sample every time.
What NVIDIA says (2)
“This function will always produce the same sample given an identical random_state .”
“frac float, optional Fraction of axis items to return.”
Random projection multiplies data by a random matrix to get fewer dimensions. The Johnson-Lindenstrauss lemma bounds how many dimensions keep distances roughly intact. PCA means principal component analysis.
What NVIDIA says (2)
“Reduce dimensionality through Gaussian random projection.”
“n_components can be automatically adjusted according to the number of samples in the dataset and the bound given by the Johnson-Lindenstrauss lemma.”
Key terms: Dimensionality reduction Principal component analysis UMAP Sampling
1.7 Efficient storage with Parquet and modern frameworks
Column pruning, filter pushdown, partitioned writes, GPUDirect Storage and the Polars GPU engine.
Key points
Parquet is a columnar file format: each column is stored separately. So a reader can load only the columns you ask for. CSV means comma-separated values. I/O means input/output.
What NVIDIA says (1)
“columns list, default None If not None, only these columns will be read.”
A row group is a block of rows inside a Parquet file. Each row group stores min/max statistics, so the reader can skip blocks that cannot match.
What NVIDIA says (1)
“If not None, specifies a filter predicate used to filter out row groups using statistics stored for each row group as Parquet metadata.”
Partitioning a dataset by a column puts each value's rows in its own folder. Readers can then skip folders they do not need.
What NVIDIA says (1)
“partition_cols list, optional, default None Column names by which to partition the dataset Columns are partitioned in the order they are given”
GPUDirect Storage lets data flow between storage and GPU memory without an extra copy through CPU memory. cuDF uses it for readers such as read_parquet when available. CSV means comma-separated values.
What NVIDIA says (1)
“cuDF leverages the KvikIO library for high-performance I/O features, such as parallel I/O operations and NVIDIA Magnum IO GPUDirect Storage (GDS).”
Polars is a DataFrame library with a lazy API, which builds a query plan before running it. cuDF provides a GPU engine for that plan. API means application programming interface.
What NVIDIA says (2)
“cuDF provides GPU-accelerated execution engines for Python users of the Polars Lazy API.”
“If it is not, the execution transparently falls back to the standard Polars engine and runs on the CPU.”
Key terms: Polars GPU engine KvikIO Parquet Partitioned dataset