Data Manipulation and Preparation

23% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

1.1 Joining and manipulating data with cuDF and pandas

Official objective: “Data integration, joining, and manipulation using NVIDIA cuDF and pandas”

What cuDF is, cudf.pandas zero-code-change acceleration, joins, group-bys, and why row loops are slow on GPUs.

Key points

  1. A DataFrame is a table with named columns. cuDF holds DataFrames in GPU memory and offers the same style of API as pandas, the most common Python table library. API means application programming interface. CUDA means NVIDIA's parallel computing platform.

    What NVIDIA says (1)

    “cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”

    — cuDF: 10 Minutes to cuDF and Dask cuDF

  2. A join (merge) combines two tables on matching key values. pandas keeps or sorts the key order; cuDF does not, because ordering costs time on a parallel GPU. Sort by the key or index after the merge when order matters.

    What NVIDIA says (2)

    “By contrast, cuDF’s default behavior is to return rows in a non-deterministic order to maximize performance.”

    — cuDF: Comparison of cuDF and pandas

    “Note that the dataframe order is not maintained , but may be restored post-merge by sorting by the index.”

    — cuDF: 10 Minutes to cuDF and Dask cuDF

  3. cudf.pandas is cuDF's pandas accelerator mode. You keep your pandas code; supported operations run on the GPU, and anything else falls back to normal pandas. API means application programming interface. SQL means Structured Query Language.

    What NVIDIA says (2)

    “It supports 100% of the Pandas API , using the GPU for supported operations, and automatically falling back to pandas for other operations.”

    — cuDF: cudf.pandas

    “Nothing changes, not even your import statements, when going from CPU to GPU.”

    — cuDF: cudf.pandas

  4. cudf.pandas wraps each object in a proxy, a stand-in that can point to a GPU or a CPU copy. It tries the GPU first and falls back to the CPU only for the failing step.

    What NVIDIA says (2)

    “Attribute lookups and method calls are first attempted on the GPU (copying from CPU if necessary).”

    — cuDF: How cudf.pandas works

    “If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”

    — cuDF: How cudf.pandas works

  5. GPUs are built to process many values at once, not one at a time. A vectorized function works on a whole column in one call.

    What NVIDIA says (2)

    “This is because iterating over data that resides on the GPU will yield extremely poor performance, as GPUs are optimized for highly parallel operations rather than sequential operations.”

    — cuDF: Comparison of cuDF and pandas

    “If you absolutely must iterate, copy the data from GPU to CPU by using .to_arrow() or .to_pandas() , then convert the result back to GPU using a Series , DataFrame or Index constructor.”

    — cuDF: Comparison of cuDF and pandas

  6. cuDF mirrors common pandas calls such as read_csv, groupby and agg. For code like this, swapping the import is enough. API means application programming interface. CUDA means NVIDIA's parallel computing platform.

    What NVIDIA says (1)

    “With RAPIDS, all you have to do to use this code to run on a GPU and enjoy the interactive querying of data is to change the import statement.”

    — Beginner's Guide to GPU-Accelerated DataFrames for pandas Users

Key terms: pandas cuDF cudf.pandas CPU fallback Join (merge) Group-by aggregation

Practice 1.1 (6 questions) Objective page

1.2 Cleaning data and handling quality and governance

Official objective: “Data cleaning, quality handling, and governance compliance”

Nulls in cuDF, dropna and fillna, duplicates, and protecting personal data.

Key points

  1. A missing value is a cell with no data, also called a null. cuDF allows nulls in every data type and shows them as <NA>. NA means not available (missing).

    What NVIDIA says (2)

    “cudf supports having missing values in all dtypes.”

    — cuDF: Working with missing data

    “To detect missing values, you can use isna() and notna() functions.”

    — cuDF: Working with missing data

  2. dropna removes rows or columns with nulls. how='any' drops a row with at least one null; how='all' drops only fully empty rows.

    What NVIDIA says (1)

    “any (default) drops rows (or columns) containing at least one null value. all drops only rows (or columns) containing all null values.”

    — cuDF API: DataFrame.dropna

  3. Imputation means filling missing values with a chosen value. fillna takes a scalar, a Series, or a dict of per-column values.

    What NVIDIA says (1)

    “A dict can be used to provide different values to fill nulls in different columns.”

    — cuDF API: DataFrame.fillna

  4. Data governance means rules about who may use which data, and how. Compliance means following those rules and the law, such as privacy rules for personal data.

    What NVIDIA says (2)

    “In addition to complying with privacy and consumer protection laws, trustworthy AI models are tested for safety, security and mitigation of unwanted bias.”

    — What Is Trustworthy AI?

    “Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”

    — Enhancing Generative AI Model Accuracy with NVIDIA NeMo Curator

  5. Data quality means data is correct, complete and free of junk such as duplicates. drop_duplicates() keeps the first copy by default and drops the rest.

    What NVIDIA says (2)

    “Training models on these datasets without proper processing can result in higher training time and lower model quality.”

    — Mastering LLM Techniques: Text Data Processing

    “Determines which duplicates (if any) to keep. - ‘first’ : Drop duplicates except for the first occurrence.”

    — cuDF API: DataFrame.drop_duplicates

Key terms: Null (missing value) dropna fillna Deduplication Personal identifiable information Data governance

Practice 1.2 (5 questions) Objective page

1.3 GPU-accelerated ETL with RAPIDS, Dask and Spark

Official objective: “GPU-accelerated ETL workflows with RAPIDS, Dask, or Spark”

What ETL is, Dask cuDF for data larger than one GPU, and the RAPIDS Accelerator for Apache Spark.

Key points

  1. ETL is the data preparation stage before analysis or training. RAPIDS accelerates it on GPUs.

    What NVIDIA says (2)

    “These tasks are often grouped under the term Extract, Transform, Load (ETL).”

    — RAPIDS Accelerates Data Science End-to-End

    “RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

    — RAPIDS Accelerates Data Science End-to-End

  2. Dask is a Python library for parallel computing. Dask cuDF registers cuDF as the DataFrame backend, so each partition is a cuDF DataFrame. To use several GPUs you also need a dask.distributed cluster; Dask-CUDA makes that easy. CSV means comma-separated values.

    What NVIDIA says (3)

    “When installed, Dask cuDF is automatically registered as the "cudf" dataframe backend for Dask DataFrame .”

    — Dask cuDF documentation

    “You must also deploy a dask.distributed cluster to leverage multiple GPUs.”

    — Dask cuDF documentation

    “Therefore to solve this with the libraries we also need to use Dask to partition the dataset and stream it through GPU memory, then cuDF can process each partition in a performant way.”

    — RAPIDS Deployment: One Billion Row Challenge on a single node

  3. Apache Spark is a distributed engine for large-scale data processing. The RAPIDS Accelerator plugs into Spark and runs supported operations on GPUs. ETL means extract, transform, load. SQL means Structured Query Language.

    What NVIDIA says (2)

    “Run your existing Apache Spark applications with no code change.”

    — cuDF for Apache Spark: Overview

    “The cuDF for Apache Spark combines the power of the RAPIDS cuDF library and the scale of the Spark distributed computing framework.”

    — cuDF for Apache Spark: Overview

  4. Feature engineering means building model inputs from raw data. AT&T's case shows GPUs can pay off in early pipeline stages too, not only in training. ETL means extract, transform, load.

    What NVIDIA says (2)

    “The optimization came from using the RAPIDS accelerator for Apache Spark, an open-source library that enables GPU-accelerated ETL and feature engineering.”

    — Scaling Data Pipelines: AT&T Optimizes Speed, Cost, and Efficiency with GPUs

    “SPOILER ALERT: We were pleasantly surprised that, at least for the examples examined, the use of GPUs for each pipeline stage proved to be faster, cheaper, and simpler!”

    — Scaling Data Pipelines: AT&T Optimizes Speed, Cost, and Efficiency with GPUs

  5. Spilling means moving data from GPU memory to host (CPU) memory to free space. Dask-CUDA has several spilling options; cuDF-level spilling often works best for tables. ETL means extract, transform, load.

    What NVIDIA says (1)

    “We’ve found that for tabular-based workloads, using enable-cudf-spill is often faster and more stable compared with the other Dask-CUDA options, including --device-memory-limit .”

    — Best Practices for Multi-GPU Data Analysis Using RAPIDS with Dask

Key terms: Dask Dask cuDF RAPIDS Accelerator for Apache Spark ETL

Practice 1.3 (5 questions) Objective page

1.4 Feature engineering for numbers and categories

Official objective: “Feature engineering for numerical and categorical variables”

One-hot and target encoding, standard scaling and binning.

Key points

  1. One-hot encoding turns a categorical column into one true/false column per category. Categorical means the values are labels, not numbers.

    What NVIDIA says (1)

    “cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”

    — cuDF API: get_dummies

  2. Target encoding replaces a category with a statistic of the label, usually its mean. cuML's TargetEncoder uses folds to avoid label leakage, which is when the label sneaks into a feature.

    What NVIDIA says (2)

    “The input data is grouped by the columns Xs and the aggregated mean value of Y of each group is calculated to replace each value of Xs .”

    — cuML API: TargetEncoder

    “Several optimizations are applied to prevent label leakage and parallelize the execution.”

    — cuML API: TargetEncoder

  3. Scaling puts numeric features on a similar range so no feature dominates by its units. StandardScaler computes z = (x - mean) / std and reuses those training statistics later.

    What NVIDIA says (2)

    “Standardize features by removing the mean and scaling to unit variance”

    — cuML API: StandardScaler

    “Mean and standard deviation are then stored to be used on later data using transform() .”

    — cuML API: StandardScaler

  4. Binning, also called discretization, groups continuous numbers into intervals. Each bin can then be treated as a category.

    What NVIDIA says (1)

    “Bin continuous data into intervals.”

    — cuML API: KBinsDiscretizer

  5. With all categories present, one column can be computed from the others. Dropping one removes that redundancy but changes how a penalized model treats categories.

    What NVIDIA says (1)

    “However, dropping one category breaks the symmetry of the original representation and can therefore induce a bias in downstream models, for instance for penalized linear classification or regression models.”

    — cuML API: OneHotEncoder

Key terms: Feature Feature engineering One-hot encoding Target encoding Standard scaling Binning

Try it: Feature lab

Practice 1.4 (5 questions) Objective page

1.5 Class imbalance and synthetic data

Official objective: “Handling class imbalance and generating synthetic data”

Why accuracy misleads, class weights, stratified splits and synthetic data.

Key points

  1. Class imbalance means one label is much rarer than another. A balanced class weight makes errors on the rare class cost more during training.

    What NVIDIA says (2)

    “The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”

    — cuML API: LogisticRegression

    “If 'balanced' , class weights are computed from the training labels.”

    — cuML API: RandomForestClassifier

  2. A stratified split keeps each class's share the same in every part. Without it, a small test set may contain few or no rare cases.

    What NVIDIA says (1)

    “If not None, data is split in a stratified fashion, using this as the class labels.”

    — cuML API: train_test_split

  3. Synthetic data is artificial data made by rules, simulations or generative models. It can fill gaps in rare classes and protect privacy.

    What NVIDIA says (2)

    “Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”

    — NVIDIA Glossary: Synthetic Data Generation

    “Data Privacy : Synthetic data helps overcome privacy issues by generating training data that mimic real-world statistics without directly corresponding to individual records.”

    — NVIDIA Glossary: Synthetic Data Generation

  4. Accuracy is the share of correct predictions. On imbalanced data it hides poor results on the rare class; a confusion matrix shows them.

    What NVIDIA says (2)

    “Accuracy classification score.”

    — cuML API: accuracy_score

    “Compute confusion matrix to evaluate the accuracy of a classification.”

    — cuML API: confusion_matrix

Key terms: Class imbalance Stratified split Synthetic data Accuracy

Try it: Confusion matrix lab

Practice 1.5 (4 questions) Objective page

1.6 Dimensionality reduction and sampling

Official objective: “Dimensionality reduction and data sampling”

PCA, UMAP, random projection and reproducible sampling.

Key points

  1. Dimensionality reduction means representing data with fewer columns while keeping most of the information. PCA keeps the directions with the most variance. PCA means principal component analysis.

    What NVIDIA says (1)

    “PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”

    — cuML API: PCA

  2. UMAP builds a nearest-neighbor graph and embeds the data into fewer dimensions. It keeps local neighborhoods, which makes clusters visible. UMAP means Uniform Manifold Approximation and Projection.

    What NVIDIA says (2)

    “UMAP is a popular dimension reduction algorithm used in fields like bioinformatics, NLP topic modeling, and ML preprocessing.”

    — Even Faster and More Scalable UMAP on the GPU with RAPIDS cuML

    “It works by creating a k-nearest neighbors (k-NN) graph, which is known in literature as an all-neighbors graph, to build a fuzzy topological representation of the data, which is used to embed high-dimensional data into lower dimensions.”

    — Even Faster and More Scalable UMAP on the GPU with RAPIDS cuML

  3. Sampling picks a subset of rows to work with. A fixed random_state gives the same sample every time.

    What NVIDIA says (2)

    “This function will always produce the same sample given an identical random_state .”

    — cuDF API: DataFrame.sample

    “frac float, optional Fraction of axis items to return.”

    — cuDF API: DataFrame.sample

  4. Random projection multiplies data by a random matrix to get fewer dimensions. The Johnson-Lindenstrauss lemma bounds how many dimensions keep distances roughly intact. PCA means principal component analysis.

    What NVIDIA says (2)

    “Reduce dimensionality through Gaussian random projection.”

    — cuML API: GaussianRandomProjection

    “n_components can be automatically adjusted according to the number of samples in the dataset and the bound given by the Johnson-Lindenstrauss lemma.”

    — cuML API: GaussianRandomProjection

Key terms: Dimensionality reduction Principal component analysis UMAP Sampling

Practice 1.6 (4 questions) Objective page

1.7 Efficient storage with Parquet and modern frameworks

Official objective: “Efficient processing and storage with Parquet and modern frameworks”

Column pruning, filter pushdown, partitioned writes, GPUDirect Storage and the Polars GPU engine.

Key points

  1. Parquet is a columnar file format: each column is stored separately. So a reader can load only the columns you ask for. CSV means comma-separated values. I/O means input/output.

    What NVIDIA says (1)

    “columns list, default None If not None, only these columns will be read.”

    — cuDF API: read_parquet

  2. A row group is a block of rows inside a Parquet file. Each row group stores min/max statistics, so the reader can skip blocks that cannot match.

    What NVIDIA says (1)

    “If not None, specifies a filter predicate used to filter out row groups using statistics stored for each row group as Parquet metadata.”

    — cuDF API: read_parquet

  3. Partitioning a dataset by a column puts each value's rows in its own folder. Readers can then skip folders they do not need.

    What NVIDIA says (1)

    “partition_cols list, optional, default None Column names by which to partition the dataset Columns are partitioned in the order they are given”

    — cuDF API: DataFrame.to_parquet

  4. GPUDirect Storage lets data flow between storage and GPU memory without an extra copy through CPU memory. cuDF uses it for readers such as read_parquet when available. CSV means comma-separated values.

    What NVIDIA says (1)

    “cuDF leverages the KvikIO library for high-performance I/O features, such as parallel I/O operations and NVIDIA Magnum IO GPUDirect Storage (GDS).”

    — cuDF: Input / Output

  5. Polars is a DataFrame library with a lazy API, which builds a query plan before running it. cuDF provides a GPU engine for that plan. API means application programming interface.

    What NVIDIA says (2)

    “cuDF provides GPU-accelerated execution engines for Python users of the Polars Lazy API.”

    — cuDF: Polars GPU engine

    “If it is not, the execution transparently falls back to the standard Polars engine and runs on the CPU.”

    — cuDF: Polars GPU engine

Key terms: Polars GPU engine KvikIO Parquet Partitioned dataset

Practice 1.7 (5 questions) Objective page