Foundations of Accelerated Data Science
12% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management
5.1 Python for data analysis
NumPy arrays, pandas DataFrames and Jupyter with cudf.pandas.
Key points
An array is a grid of values of one type. NumPy is the base that pandas and scikit-learn build on. SQL means Structured Query Language.
What NVIDIA says (2)
“What is NumPy NumPy is a powerful, well-optimized, free open-source library for the Python programming language, adding support for large, multi-dimensional arrays (also called matrices or tensors).”
“As the core library for scientific computing, NumPy is the base for libraries such as Pandas, Scikit-learn , and SciPy .”
The DataFrame is pandas' core data structure. Columns can hold numbers, categories or text. CSV means comma-separated values.
What NVIDIA says (1)
“A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable, and each row contains a set of values corresponding to those variables.”
Jupyter is an interactive notebook for code, results and notes. Magic commands like %load_ext change how the notebook runs code.
What NVIDIA says (1)
“Just %load_ext cudf.pandas in Jupyter, or pass -m cudf.pandas on the command line.”
Key terms: NumPy pandas DataFrame Jupyter notebook cudf.pandas
5.2 Why GPUs speed up data science
Accelerated computing, parallelism and measuring speedups.
Key points
Parallel processing means doing many operations at the same time. GPUs run thousands of operations at once, which suits data science on large tables.
What NVIDIA says (1)
“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”
A GPU has many cores built for throughput. A column operation applies the same step to every value, which maps well to many cores.
What NVIDIA says (1)
“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”
Benchmarking means measuring run time under the same conditions. The profiler shows fallbacks that can eat the speedup.
What NVIDIA says (2)
“Additionally, you can use Python magic commands like %%time and %%timeit to enable benchmarks of specific code blocks that facilitate direct comparisons of runtime between pandas (CPU) and the cuDF accelerator for pandas (GPU).”
“The profiler will indicate which functions ran on GPU / CPU.”
Key terms: Graphics processing unit Accelerated computing Parallel processing Benchmarking
Try it: Transfer lab
5.3 CPU versus GPU workloads and memory transfers
When small data is faster on the CPU, minimizing host-device copies and pinned memory.
Key points
Moving data between CPU and GPU memory takes time. With little data, that fixed cost dominates. GPUs shine on many rows.
What NVIDIA says (2)
“Small data sizes may be slower on GPU than CPU, because of the cost of data transfers.”
“cuDF achieves the highest performance with many rows of data.”
The host is the CPU side; the device is the GPU side. The link between them is much slower than GPU memory, so keep data on the GPU.
What NVIDIA says (2)
“Minimize the amount of data transferred between host and device when possible, even if that means running kernels on the GPU that get little or no speed-up compared to running them on the host CPU.”
“Batching many small transfers into one larger transfer performs much better because it eliminates most of the per-transfer overhead.”
Pageable memory can be moved by the operating system; pinned memory cannot. Copying from pinned memory skips an extra staging copy.
What NVIDIA says (2)
“Higher bandwidth is possible between the host and the device when using page-locked (or “pinned”) memory.”
“The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”
Key terms: Graphics processing unit Central processing unit Host and device Pinned memory
Try it: Transfer lab
5.4 The end-to-end data science workflow
Ingest, ETL, clean and transform, and keeping every stage on the GPU.
Key points
Ingest means bringing raw data in. ETL, cleaning and transformation prepare it for training. ETL means extract, transform, load.
What NVIDIA says (1)
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
Cleaning fixes types, gaps and errors so data can be analyzed. cuDF offers the same style of API on GPUs. MLOps means machine learning operations. API means application programming interface.
What NVIDIA says (1)
“With its support for structured data formats like tables, matrices, and time series, the pandas Python API provides tools to process messy or raw datasets into clean, structured formats ready for analysis.”
Each move between tools or devices costs time. Keeping the whole pipeline on the GPU removes most of those moves. ETL means extract, transform, load.
What NVIDIA says (1)
“To get the highest performance for your ML pipeline on NVIDIA GPUs, minimize data transfers between the CPUs and GPUs as part of your pipeline.”
Key terms: RAPIDS ETL Data science pipeline
5.5 Distributed versus GPU-accelerated frameworks
What Dask adds, how it compares with Spark, and Dask-CUDA for multi-GPU.
Key points
A distributed framework spreads work over many machines or processes. Dask scales NumPy, pandas and scikit-learn style code.
What NVIDIA says (2)
“Dask is an open-source library designed to provide parallelism to the existing Python stack.”
“It provides integrations with Python libraries like NumPy Arrays, Pandas DataFrames, and scikit-learn to enable parallel execution across multiple cores, processors, and computers without having to learn new libraries or languages.”
Spark has its own API and engine. RAPIDS keeps pandas-like and scikit-learn-like APIs and uses Dask to scale out. API means application programming interface.
What NVIDIA says (1)
“However, unlike Apache Spark, it does not introduce a new API but provides a familiar programming interface of tools found in the PyData ecosystem, like pandas, scikit-learn, NetworkX, and more.”
Dask cuDF is the GPU backend for Dask DataFrames. Running across GPUs or nodes needs a cluster of workers, which Dask-CUDA sets up.
What NVIDIA says (2)
“Note Neither Dask cuDF nor Dask DataFrame provide support for multi-GPU or multi-node execution on their own.”
“You must also deploy a dask.distributed cluster to leverage multiple GPUs.”
Key terms: Parallel processing Dask Dask cuDF
5.6 Parameters, tuning and fitting
Parameters versus hyperparameters, underfitting and overfitting.
Key points
Training adjusts parameters, such as tree splits or linear weights. Tuning searches hyperparameters by training many times.
What NVIDIA says (1)
“This is particularly important because data scientists typically run the algorithm not just once, but many times in order to tune hyperparameters (such as learning rate or tree depth) and find the best accuracy.”
Underfitting is high bias: the model misses real structure. Overfitting is the opposite: it learns noise and fails on new data.
What NVIDIA says (1)
“Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”
Each tree overfits in its own way. Averaging their votes cancels much of that noise.
What NVIDIA says (1)
“The presence of a large number of trees also reduces the problem of overfitting, which occurs when a model incorporates too much “noise” in the training data and makes poor decisions as a result.”
Key terms: Overfitting Underfitting Hyperparameter