Advance Data Structures

7% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

7.1 Time series splits and forecast evaluation

Official objective: “Time-series data handling, splitting, and forecasting evaluation”

Time-ordered validation, recursive and direct forecasting, lag and rolling features.

Key points

  1. Cross-validation means testing a model on several held-out slices of data. With time series, random shuffling lets the model peek at the future.

    What NVIDIA says (2)

    “Great care must be taken when defining cross-validation folds for time-series data.”

    — RAPIDS Deployment: Time series forecasting with HPO

    “We are not allowed to use the future to predict the past, so the training set must precede (in time) the validation set.”

    — RAPIDS Deployment: Time series forecasting with HPO

  2. A forecast horizon is how many future steps you predict. Direct forecasting needs many models, so it costs more compute.

    What NVIDIA says (2)

    “Multistep forecasting One popular technique used in time series forecasting is recursive multi-step forecasting , in which you train a single model and apply it recursively to predict the next n values in the series.”

    — Accelerating Time Series Forecasting with RAPIDS cuML

    “In contrast, direct multi-step forecasting uses a separate model to predict each future value in your forecast horizon.”

    — Accelerating Time Series Forecasting with RAPIDS cuML

  3. A lag feature is a past value of the target, such as last week's sales. A rolling-window statistic is a summary, such as a mean, over a recent time span.

    What NVIDIA says (2)

    “Lag features are useful because what happens in the past often influences what would happen in the future.”

    — RAPIDS Deployment: Time series forecasting with HPO

    “Rolling window statistics are statistics (e.g. mean, standard deviation) over a time duration in the past.”

    — RAPIDS Deployment: Time series forecasting with HPO

Key terms: Rolling mean Time series Forecast horizon Lag feature

Try it: Split lab

Practice 7.1 (3 questions) Objective page

7.2 Missing and irregular timestamps in cuDF

Official objective: “Managing missing or irregular timestamps with cuDF interpolation”

interpolate() and resample().

Key points

  1. Interpolation estimates missing values from their neighbors. Linear interpolation assumes a straight line between the known points. With irregular timestamps, method='index' uses the time index as the x-axis. NaN means not a number.

    What NVIDIA says (3)

    “Parameters : method str, default ‘linear’ Interpolation technique to use.”

    — cuDF API: DataFrame.interpolate

    “Returns the same object type as the caller, interpolated at some or all NaN values”

    — cuDF API: DataFrame.interpolate

    “‘index’, ‘values’: linearly interpolate using the index as an x-axis.”

    — cuDF API: DataFrame.interpolate

  2. Resampling regroups time series data onto a new, fixed time grid. You then aggregate (or interpolate) within each bin.

    What NVIDIA says (2)

    “Parameters : rule: str The offset string representing the frequency to use.”

    — cuDF API: DataFrame.resample

    “First, we create a time series with 1 minute intervals”

    — cuDF API: DataFrame.resample

Key terms: Interpolation Resampling

Try it: Interpolation lab

Practice 7.2 (2 questions) Objective page

7.3 CPU versus GPU for temporal analytics

Official objective: “CPU vs. GPU performance for temporal analytics”

Why forecasting gets expensive and how cuML drops into skforecast.

Key points

  1. Each extra model and each extra step multiplies the compute. GPUs run these many similar computations in parallel.

    What NVIDIA says (1)

    “With growing datasets and techniques like direct multi-step forecasting that require you to run several models at once, forecasts can quickly become computationally expensive when running on CPU-based infrastructure.”

    — Accelerating Time Series Forecasting with RAPIDS cuML

  2. cuML has a scikit-learn compatible API. Tools built on that API can use cuML models as drop-in replacements. API means application programming interface.

    What NVIDIA says (2)

    “Bringing accelerated computing to direct multistep forecasting cuML can be dropped into existing skforecast workflows.”

    — Accelerating Time Series Forecasting with RAPIDS cuML

    “NVIDIA cuML is a GPU-accelerated machine learning library for Python with a scikit-learn compatible API.”

    — Accelerating Time Series Forecasting with RAPIDS cuML

Key terms: Central processing unit

Practice 7.3 (2 questions) Objective page

7.4 Representing data as graphs

Official objective: “Graph-based data representation and analysis”

Nodes and edges, and building a cuGraph graph from an edge list.

Key points

  1. A graph is a structure for modeling relationships. Graph analytics studies pairwise relationships and the structure of the whole network.

    What NVIDIA says (1)

    “A graph consists of nodes or vertices (representing the entities in the system) that are connected by edges (representing relationships between those entities).”

    — NVIDIA Glossary: Graph Analytics

  2. An edge list is a table with one row per edge. cuGraph builds a GPU graph object directly from it. CSV means comma-separated values.

    What NVIDIA says (2)

    “This cudf.DataFrame contains columns storing edge source vertices, destination (or target following NetworkX’s terminology) vertices”

    — cuGraph API: from_cudf_edgelist

    “A GPU Graph Object (Base class of other graph types)”

    — cuGraph API: Graph

Key terms: cuGraph Graph Edge list

Try it: PageRank lab

Practice 7.4 (2 questions) Objective page

7.5 Node importance and network visualization

Official objective: “Node importance evaluation and network relationship visualization”

PageRank scores and ForceAtlas2 layouts.

Key points

  1. PageRank scores a node higher when important nodes link to it. It is a classic measure of node importance (centrality).

    What NVIDIA says (1)

    “Find the PageRank score for every vertex in a graph.”

    — cuGraph API: pagerank

  2. A graph layout assigns x and y positions to nodes for plotting. Force-directed layouts pull linked nodes together.

    What NVIDIA says (1)

    “ForceAtlas2 is a continuous graph layout algorithm for handy network visualization.”

    — cuGraph API: force_atlas2

Key terms: cuGraph PageRank Graph layout

Try it: PageRank lab

Practice 7.5 (2 questions) Objective page