Foundations of Accelerated Data Science

12% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

5.1 Python for data analysis

Official objective: “Python fundamentals for data analysis (NumPy, pandas, Jupyter)”

NumPy arrays, pandas DataFrames and Jupyter with cudf.pandas.

Key points

  1. An array is a grid of values of one type. NumPy is the base that pandas and scikit-learn build on. SQL means Structured Query Language.

    What NVIDIA says (2)

    “What is NumPy NumPy is a powerful, well-optimized, free open-source library for the Python programming language, adding support for large, multi-dimensional arrays (also called matrices or tensors).”

    — NVIDIA Glossary: NumPy

    “As the core library for scientific computing, NumPy is the base for libraries such as Pandas, Scikit-learn , and SciPy .”

    — NVIDIA Glossary: NumPy

  2. The DataFrame is pandas' core data structure. Columns can hold numbers, categories or text. CSV means comma-separated values.

    What NVIDIA says (1)

    “A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable, and each row contains a set of values corresponding to those variables.”

    — What Is Pandas and Why Does it Matter?

  3. Jupyter is an interactive notebook for code, results and notes. Magic commands like %load_ext change how the notebook runs code.

    What NVIDIA says (1)

    “Just %load_ext cudf.pandas in Jupyter, or pass -m cudf.pandas on the command line.”

    — cuDF: cudf.pandas

Key terms: NumPy pandas DataFrame Jupyter notebook cudf.pandas

Practice 5.1 (3 questions) Objective page

5.2 Why GPUs speed up data science

Official objective: “Core GPU acceleration concepts and advantages for data science”

Accelerated computing, parallelism and measuring speedups.

Key points

  1. Parallel processing means doing many operations at the same time. GPUs run thousands of operations at once, which suits data science on large tables.

    What NVIDIA says (1)

    “Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”

    — What Is Accelerated Computing?

  2. A GPU has many cores built for throughput. A column operation applies the same step to every value, which maps well to many cores.

    What NVIDIA says (1)

    “By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”

    — What's the Difference Between a CPU and a GPU?

  3. Benchmarking means measuring run time under the same conditions. The profiler shows fallbacks that can eat the speedup.

    What NVIDIA says (2)

    “Additionally, you can use Python magic commands like %%time and %%timeit to enable benchmarks of specific code blocks that facilitate direct comparisons of runtime between pandas (CPU) and the cuDF accelerator for pandas (GPU).”

    — Get Started with GPU Acceleration for Data Science

    “The profiler will indicate which functions ran on GPU / CPU.”

    — cuDF: cudf.pandas FAQ and Known Issues

Key terms: Graphics processing unit Accelerated computing Parallel processing Benchmarking

Try it: Transfer lab

Practice 5.2 (3 questions) Objective page

5.3 CPU versus GPU workloads and memory transfers

Official objective: “CPU vs. GPU workloads and memory transfer optimization”

When small data is faster on the CPU, minimizing host-device copies and pinned memory.

Key points

  1. Moving data between CPU and GPU memory takes time. With little data, that fixed cost dominates. GPUs shine on many rows.

    What NVIDIA says (2)

    “Small data sizes may be slower on GPU than CPU, because of the cost of data transfers.”

    — cuDF: cudf.pandas FAQ and Known Issues

    “cuDF achieves the highest performance with many rows of data.”

    — cuDF: cudf.pandas FAQ and Known Issues

  2. The host is the CPU side; the device is the GPU side. The link between them is much slower than GPU memory, so keep data on the GPU.

    What NVIDIA says (2)

    “Minimize the amount of data transferred between host and device when possible, even if that means running kernels on the GPU that get little or no speed-up compared to running them on the host CPU.”

    — How to Optimize Data Transfers in CUDA C/C++

    “Batching many small transfers into one larger transfer performs much better because it eliminates most of the per-transfer overhead.”

    — How to Optimize Data Transfers in CUDA C/C++

  3. Pageable memory can be moved by the operating system; pinned memory cannot. Copying from pinned memory skips an extra staging copy.

    What NVIDIA says (2)

    “Higher bandwidth is possible between the host and the device when using page-locked (or “pinned”) memory.”

    — How to Optimize Data Transfers in CUDA C/C++

    “The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”

    — How to Optimize Data Transfers in CUDA C/C++

Key terms: Graphics processing unit Central processing unit Host and device Pinned memory

Try it: Transfer lab

Practice 5.3 (3 questions) Objective page

5.4 The end-to-end data science workflow

Official objective: “End-to-end data science workflow (ingest, ETL, clean, transform)”

Ingest, ETL, clean and transform, and keeping every stage on the GPU.

Key points

  1. Ingest means bringing raw data in. ETL, cleaning and transformation prepare it for training. ETL means extract, transform, load.

    What NVIDIA says (1)

    “RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

    — RAPIDS Accelerates Data Science End-to-End

  2. Cleaning fixes types, gaps and errors so data can be analyzed. cuDF offers the same style of API on GPUs. MLOps means machine learning operations. API means application programming interface.

    What NVIDIA says (1)

    “With its support for structured data formats like tables, matrices, and time series, the pandas Python API provides tools to process messy or raw datasets into clean, structured formats ready for analysis.”

    — What Is Pandas and Why Does it Matter?

  3. Each move between tools or devices costs time. Keeping the whole pipeline on the GPU removes most of those moves. ETL means extract, transform, load.

    What NVIDIA says (1)

    “To get the highest performance for your ML pipeline on NVIDIA GPUs, minimize data transfers between the CPUs and GPUs as part of your pipeline.”

    — NVIDIA cuML Brings Zero Code Change Acceleration to scikit-learn

Key terms: RAPIDS ETL Data science pipeline

Practice 5.4 (3 questions) Objective page

5.5 Distributed versus GPU-accelerated frameworks

Official objective: “Distributed vs. GPU-accelerated computing frameworks”

What Dask adds, how it compares with Spark, and Dask-CUDA for multi-GPU.

Key points

  1. A distributed framework spreads work over many machines or processes. Dask scales NumPy, pandas and scikit-learn style code.

    What NVIDIA says (2)

    “Dask is an open-source library designed to provide parallelism to the existing Python stack.”

    — NVIDIA Glossary: Dask

    “It provides integrations with Python libraries like NumPy Arrays, Pandas DataFrames, and scikit-learn to enable parallel execution across multiple cores, processors, and computers without having to learn new libraries or languages.”

    — NVIDIA Glossary: Dask

  2. Spark has its own API and engine. RAPIDS keeps pandas-like and scikit-learn-like APIs and uses Dask to scale out. API means application programming interface.

    What NVIDIA says (1)

    “However, unlike Apache Spark, it does not introduce a new API but provides a familiar programming interface of tools found in the PyData ecosystem, like pandas, scikit-learn, NetworkX, and more.”

    — Beginner's Guide to GPU-Accelerated DataFrames for pandas Users

  3. Dask cuDF is the GPU backend for Dask DataFrames. Running across GPUs or nodes needs a cluster of workers, which Dask-CUDA sets up.

    What NVIDIA says (2)

    “Note Neither Dask cuDF nor Dask DataFrame provide support for multi-GPU or multi-node execution on their own.”

    — Dask cuDF documentation

    “You must also deploy a dask.distributed cluster to leverage multiple GPUs.”

    — Dask cuDF documentation

Key terms: Parallel processing Dask Dask cuDF

Practice 5.5 (3 questions) Objective page

5.6 Parameters, tuning and fitting

Official objective: “Model parameters, tuning, and overfitting vs. underfitting concepts”

Parameters versus hyperparameters, underfitting and overfitting.

Key points

  1. Training adjusts parameters, such as tree splits or linear weights. Tuning searches hyperparameters by training many times.

    What NVIDIA says (1)

    “This is particularly important because data scientists typically run the algorithm not just once, but many times in order to tune hyperparameters (such as learning rate or tree depth) and find the best accuracy.”

    — Gradient Boosting, Decision Trees and XGBoost with CUDA

  2. Underfitting is high bias: the model misses real structure. Overfitting is the opposite: it learns noise and fails on new data.

    What NVIDIA says (1)

    “Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”

    — What Is XGBoost and Why Does It Matter?

  3. Each tree overfits in its own way. Averaging their votes cancels much of that noise.

    What NVIDIA says (1)

    “The presence of a large number of trees also reduces the problem of overfitting, which occurs when a model incorporates too much “noise” in the training data and makes poor decisions as a result.”

    — NVIDIA Glossary: Random Forest

Key terms: Overfitting Underfitting Hyperparameter

Practice 5.6 (3 questions) Objective page