Advance Data Structures
7% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management
7.1 Time series splits and forecast evaluation
Time-ordered validation, recursive and direct forecasting, lag and rolling features.
Key points
Cross-validation means testing a model on several held-out slices of data. With time series, random shuffling lets the model peek at the future.
What NVIDIA says (2)
“Great care must be taken when defining cross-validation folds for time-series data.”
“We are not allowed to use the future to predict the past, so the training set must precede (in time) the validation set.”
A forecast horizon is how many future steps you predict. Direct forecasting needs many models, so it costs more compute.
What NVIDIA says (2)
“Multistep forecasting One popular technique used in time series forecasting is recursive multi-step forecasting , in which you train a single model and apply it recursively to predict the next n values in the series.”
“In contrast, direct multi-step forecasting uses a separate model to predict each future value in your forecast horizon.”
A lag feature is a past value of the target, such as last week's sales. A rolling-window statistic is a summary, such as a mean, over a recent time span.
What NVIDIA says (2)
“Lag features are useful because what happens in the past often influences what would happen in the future.”
“Rolling window statistics are statistics (e.g. mean, standard deviation) over a time duration in the past.”
Key terms: Rolling mean Time series Forecast horizon Lag feature
Try it: Split lab
7.2 Missing and irregular timestamps in cuDF
interpolate() and resample().
Key points
Interpolation estimates missing values from their neighbors. Linear interpolation assumes a straight line between the known points. With irregular timestamps, method='index' uses the time index as the x-axis. NaN means not a number.
What NVIDIA says (3)
“Parameters : method str, default ‘linear’ Interpolation technique to use.”
“Returns the same object type as the caller, interpolated at some or all NaN values”
“‘index’, ‘values’: linearly interpolate using the index as an x-axis.”
Resampling regroups time series data onto a new, fixed time grid. You then aggregate (or interpolate) within each bin.
What NVIDIA says (2)
“Parameters : rule: str The offset string representing the frequency to use.”
“First, we create a time series with 1 minute intervals”
Key terms: Interpolation Resampling
Try it: Interpolation lab
7.3 CPU versus GPU for temporal analytics
Why forecasting gets expensive and how cuML drops into skforecast.
Key points
Each extra model and each extra step multiplies the compute. GPUs run these many similar computations in parallel.
What NVIDIA says (1)
“With growing datasets and techniques like direct multi-step forecasting that require you to run several models at once, forecasts can quickly become computationally expensive when running on CPU-based infrastructure.”
cuML has a scikit-learn compatible API. Tools built on that API can use cuML models as drop-in replacements. API means application programming interface.
What NVIDIA says (2)
“Bringing accelerated computing to direct multistep forecasting cuML can be dropped into existing skforecast workflows.”
“NVIDIA cuML is a GPU-accelerated machine learning library for Python with a scikit-learn compatible API.”
Key terms: Central processing unit
7.4 Representing data as graphs
Nodes and edges, and building a cuGraph graph from an edge list.
Key points
A graph is a structure for modeling relationships. Graph analytics studies pairwise relationships and the structure of the whole network.
What NVIDIA says (1)
“A graph consists of nodes or vertices (representing the entities in the system) that are connected by edges (representing relationships between those entities).”
An edge list is a table with one row per edge. cuGraph builds a GPU graph object directly from it. CSV means comma-separated values.
What NVIDIA says (2)
“This cudf.DataFrame contains columns storing edge source vertices, destination (or target following NetworkX’s terminology) vertices”
“A GPU Graph Object (Base class of other graph types)”
Key terms: cuGraph Graph Edge list
Try it: PageRank lab
7.5 Node importance and network visualization
PageRank scores and ForceAtlas2 layouts.
Key points
PageRank scores a node higher when important nodes link to it. It is a classic measure of node importance (centrality).
What NVIDIA says (1)
“Find the PageRank score for every vertex in a graph.”
A graph layout assigns x and y positions to nodes for plotting. Force-directed layouts pull linked nodes together.
What NVIDIA says (1)
“ForceAtlas2 is a continuous graph layout algorithm for handy network visualization.”
Key terms: cuGraph PageRank Graph layout
Try it: PageRank lab