NCA-ADS glossary

The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.

A

Accelerated computing

Using specialized hardware such as GPUs to speed up work through parallel processing.

Objectives: 5.2

What NVIDIA says (1)

“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”

— What Is Accelerated Computing?

Accuracy

The share of predictions that are correct.

Objectives: 2.6, 1.5

What NVIDIA says (2)

“Accuracy score is the ratio of correct predictions to the total number of predictions.”

— cuML: Training and evaluating machine learning models

“Accuracy classification score.”

— cuML API: accuracy_score

B

Bagging

Training many models on random samples and averaging them to reduce variance.

Objectives: 3.3

What NVIDIA says (2)

“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”

— NVIDIA Glossary: Random Forest

“Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”

— NVIDIA Glossary: Random Forest

Benchmarking

Timing the same workload under different setups, such as CPU and GPU, to compare them fairly.

Objectives: 6.6, 5.2

What NVIDIA says (2)

“In the notebook demo below, we compare benchmarking results to show how GPU can accelerate HPO tuning jobs relative to CPU.”

— RAPIDS Deployment: XGBoost and Random Forest GPU vs CPU benchmark

“Additionally, you can use Python magic commands like %%time and %%timeit to enable benchmarks of specific code blocks that facilitate direct comparisons of runtime between pandas (CPU) and the cuDF accelerator for pandas (GPU).”

— Get Started with GPU Acceleration for Data Science

Binning

Grouping a continuous number into ranges, such as age bands.

Objectives: 1.4

What NVIDIA says (1)

“Bin continuous data into intervals.”

— cuML API: KBinsDiscretizer

Boosting

Training models one after another, each fixing the errors of the last, to reduce bias.

Objectives: 3.3

What NVIDIA says (1)

“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”

— NVIDIA Glossary: Random Forest

C

Central processing unit (CPU)

A computer's general-purpose processor, with a few fast cores that suit step-by-step work and small data.

Objectives: 5.3, 7.3

What NVIDIA says (2)

“Small data sizes may be slower on GPU than CPU, because of the cost of data transfers.”

— cuDF: cudf.pandas FAQ and Known Issues

“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”

— What's the Difference Between a CPU and a GPU?

Class imbalance

When one class is much rarer than the others in the data.

Objectives: 1.5

What NVIDIA says (2)

“The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”

— cuML API: LogisticRegression

“Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”

— NVIDIA Glossary: Synthetic Data Generation

Classification

Predicting a category, such as spam or not spam.

Objectives: 2.2

What NVIDIA says (1)

“Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”

— NVIDIA Glossary: Linear Regression and Logistic Regression

Clustering

Grouping similar unlabeled rows together.

Objectives: 2.2

What NVIDIA says (2)

“Common unsupervised tasks include clustering and association.”

— NVIDIA Glossary: K-Means

“DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”

— cuML API: DBSCAN

conda

A package and environment manager that installs Python and non-Python packages, including CUDA libraries.

Objectives: 8.2, 8.3

What NVIDIA says (2)

“When installing libraries via conda, the package manager automatically pulls the required CUDA runtime libraries alongside CuPy and other dependencies, providing complete dependency management in a single installation step.”

— RAPIDS Deployment: Custom RAPIDS Docker guide

“The defaults channel is not supported by these packages, which are built to be compatible with dependencies from the conda-forge channel.”

— NVIDIA CUDA-X Data Science: Installation guide

Conda channel

A source that conda downloads packages from, such as conda-forge.

Objectives: 8.3

What NVIDIA says (1)

“The defaults channel is not supported by these packages, which are built to be compatible with dependencies from the conda-forge channel.”

— NVIDIA CUDA-X Data Science: Installation guide

Confidence interval

A range that likely contains the true value; more samples make it narrower.

Objectives: 4.4

What NVIDIA says (2)

“A single success rate on N rollouts tells you almost nothing about how confident you should be in a policy’s true performance.”

— How to Evaluate General-Purpose Robot Policies for Real-World Deployment

“Narrowing the confidence interval from 10 to 2 percentage points requires roughly 15x more rollouts (70 to 1,030).”

— How to Evaluate General-Purpose Robot Policies for Real-World Deployment

Confusion matrix

A table that counts each true class against each predicted class.

Objectives: 2.6

What NVIDIA says (2)

“Compute confusion matrix to evaluate the accuracy of a classification.”

— cuML API: confusion_matrix

“Normalizes confusion matrix over the true (rows), predicted (columns) conditions or al”

— cuML API: confusion_matrix

Container

A packaged environment with an app and its dependencies that runs the same on any host.

Objectives: 8.2, 8.1

What NVIDIA says (2)

“These size reductions result in faster container pulls and deployments, reduced storage costs in container registries, lower bandwidth usage in distributed environments, and quicker startup times for containerized applications.”

— RAPIDS Deployment: Custom RAPIDS Docker guide

“AI Workbench provides reproducibility by managing software, containers and Git repositories.”

— NVIDIA AI Workbench: Introduction

Correlation

A measure of how two variables move together; it does not prove cause.

Objectives: 4.5

What NVIDIA says (2)

“pearson : Standard correlation coefficient spearman : Spearman rank correlation”

— cuDF API: DataFrame.corr

“However, it’s important to remember that correlation and causation are two different things.”

— NVIDIA Glossary: Linear Regression and Logistic Regression

CPU fallback

Running an operation on the CPU when the GPU path does not support it, at the cost of extra data copies.

Objectives: 1.1, 2.1

What NVIDIA says (2)

“If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”

— cuDF: How cudf.pandas works

“CPU fallback preserves compatibility, but frequent transitions between CPU and GPU execution may reduce the overall speedup.”

— cuML: cuml.accel usage

Cross-filtering

Linked charts where selecting data in one filters all the others.

Objectives: 4.3

What NVIDIA says (1)

“Instead of creating several individual group by and query operations, a cuxfilter dashboard can simply cross-link numerous charts to quickly find patterns or anomalies (Figure 5).”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Cross-validation

Rotating which part of the data is held out, to get a steadier performance estimate.

Objectives: 2.5

What NVIDIA says (2)

“Cross-validation is often used to more accurately estimate the performance of the models in the search process.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

“Cross-validation is the method of splitting the training set into complementary subsets and performing training on one of the subsets, then predicting the models performance on the other.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

CUDA driver

The NVIDIA system software that lets programs use the GPU; it must support the CUDA version your packages need.

Objectives: 8.4

What NVIDIA says (2)

“You will have to ensure the CUDA driver on your machine supports the CUDA version you are trying to install with conda.”

— NVIDIA CUDA-X Data Science: Installation guide

“If conda has incorrectly identified the CUDA driver, you can override by setting the CONDA_OVERRIDE_CUDA environment variable.”

— NVIDIA CUDA-X Data Science: Installation guide

CUDA_VISIBLE_DEVICES

An environment variable that sets which GPUs a program can see.

Objectives: 8.4

What NVIDIA says (1)

“To specify a device to run on, we recommend using the CUDA_VISIBLE_DEVICES ( doc ) environment variable.”

— cuML: Advanced usage (device selection)

cuDF

A GPU DataFrame library with a pandas-style API for loading, joining, aggregating and filtering data.

Objectives: 1.1

What NVIDIA says (1)

“cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”

— cuDF: 10 Minutes to cuDF and Dask cuDF

cudf.pandas

A cuDF mode that runs existing pandas code on the GPU without code changes and falls back to the CPU when needed.

Objectives: 1.1, 5.1

What NVIDIA says (2)

“It supports 100% of the Pandas API , using the GPU for supported operations, and automatically falling back to pandas for other operations.”

— cuDF: cudf.pandas

“If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”

— cuDF: How cudf.pandas works

cuGraph

A RAPIDS library for graph analytics on the GPU.

Objectives: 7.4, 7.5

What NVIDIA says (2)

“A GPU Graph Object (Base class of other graph types)”

— cuGraph API: Graph

“Find the PageRank score for every vertex in a graph.”

— cuGraph API: pagerank

cuML

A GPU machine learning library whose estimators work like scikit-learn estimators.

Objectives: 2.1

What NVIDIA says (2)

“cuML estimators look and feel just like scikit-learn estimators .”

— cuML: Introduction

“NVIDIA cuML is a GPU-accelerated machine learning library for Python with a scikit-learn compatible API.”

— Accelerating Time Series Forecasting with RAPIDS cuML

cuml.accel

A cuML mode that runs existing scikit-learn code on the GPU without code changes.

Objectives: 2.1

What NVIDIA says (2)

“When running a script, use the cuml.accel command-line interface: python -m cuml.accel script.py In Jupyter or IPython, load the extension before other imports: % load_ext cuml.accel”

— cuML: cuml.accel usage

“Enable cuml.accel before importing scikit-learn, UMAP, or HDBSCAN.”

— cuML: cuml.accel usage

cuxfilter

A RAPIDS library for GPU dashboards whose charts filter each other.

Objectives: 4.3

What NVIDIA says (1)

“Instead of creating several individual group by and query operations, a cuxfilter dashboard can simply cross-link numerous charts to quickly find patterns or anomalies (Figure 5).”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

D

Dask

A Python library that runs familiar tools such as pandas in parallel across cores, GPUs and machines.

Objectives: 5.5, 1.3

What NVIDIA says (2)

“Dask is an open-source library designed to provide parallelism to the existing Python stack.”

— NVIDIA Glossary: Dask

“However, unlike Apache Spark, it does not introduce a new API but provides a familiar programming interface of tools found in the PyData ecosystem, like pandas, scikit-learn, NetworkX, and more.”

— Beginner's Guide to GPU-Accelerated DataFrames for pandas Users

Dask cuDF

Dask DataFrames backed by cuDF, for data split into partitions across GPUs.

Objectives: 1.3, 5.5

What NVIDIA says (2)

“When installed, Dask cuDF is automatically registered as the "cudf" dataframe backend for Dask DataFrame .”

— Dask cuDF documentation

“You must also deploy a dask.distributed cluster to leverage multiple GPUs.”

— Dask cuDF documentation

Data augmentation

Adding new rows or derived columns to make training data richer.

Objectives: 3.4

What NVIDIA says (2)

“Synthetically generated data can augment existing data for larger, more representative datasets.”

— NVIDIA Glossary: Synthetic Data Generation

“Augmenting the data using RAPIDS cuSpatial to quickly calculate distances also shows that most trips are relatively short.”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Data drift

A change in the input data between training and production.

Objectives: 6.4

What NVIDIA says (1)

“Data drift : Changes in distribution between the training data and production data can be monitored to check for drift: this is done by detecting changes in the statistical properties of feature values over time.”

— A Guide to Monitoring Machine Learning Models in Production

Data flywheel

A loop where production feedback and new data keep improving a deployed model.

Objectives: 3.5

What NVIDIA says (2)

“AI data flywheels work by creating a loop where AI models continuously improve by learning from the latest institutional knowledge and user feedback.”

— Data flywheel: What it is and how it works

“As the system interacts with the environment, it collects feedback and new data, which are then used to refine and enhance the backbone models powering the AI workflows.”

— Data flywheel: What it is and how it works

Data governance

Rules for using data lawfully and responsibly, including privacy and consent.

Objectives: 1.2

What NVIDIA says (2)

“In addition to complying with privacy and consumer protection laws, trustworthy AI models are tested for safety, security and mitigation of unwanted bias.”

— What Is Trustworthy AI?

“Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”

— Enhancing Generative AI Model Accuracy with NVIDIA NeMo Curator

Data science pipeline

The chain of steps from loading data to training and serving a model.

Objectives: 3.1, 5.4

What NVIDIA says (2)

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

“Each stage can also run independently for ad-hoc experiments.”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

DataFrame

A two-dimensional table where each column is one variable and each row is one record.

Objectives: 5.1

What NVIDIA says (1)

“A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable, and each row contains a set of values corresponding to those variables.”

— What Is Pandas and Why Does it Matter?

Datashader

A library that turns millions of points into an accurate image by aggregating them first.

Objectives: 4.2

What NVIDIA says (2)

“The Datashader library directly supports cuDF and can rapidly render over millions of aggregated points.”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

“Datapoint rendering displaying high-resolution patterns is precisely what Datashader is designed for.”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

DBSCAN (Density-Based Spatial Clustering of Applications with Noise)

A clustering method that finds dense regions of any shape and marks sparse points as noise.

Objectives: 2.2

What NVIDIA says (2)

“DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”

— cuML API: DBSCAN

“This also allows DBSCAN to be robust to noise.”

— cuML API: DBSCAN

Deduplication

Removing repeated records so they do not skew training.

Objectives: 1.2

What NVIDIA says (2)

“Training models on these datasets without proper processing can result in higher training time and lower model quality.”

— Mastering LLM Techniques: Text Data Processing

“Determines which duplicates (if any) to keep. - ‘first’ : Drop duplicates except for the first occurrence.”

— cuDF API: DataFrame.drop_duplicates

Descriptive statistics

Summary numbers such as count, mean, standard deviation and percentiles.

Objectives: 4.1

What NVIDIA says (2)

“For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”

— cuDF API: DataFrame.describe

“The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”

— cuDF API: DataFrame.describe

Dimensionality reduction

Representing data with fewer columns while keeping most of its structure.

Objectives: 1.6

What NVIDIA says (2)

“PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”

— cuML API: PCA

“Reduce dimensionality through Gaussian random projection.”

— cuML API: GaussianRandomProjection

Dockerfile

A text recipe for building a container image.

Objectives: 8.2

What NVIDIA says (1)

“To begin, you will need to create a few local files for your custom build: a Dockerfile and a configuration file ( env.yaml for conda or requirements.txt for pip).”

— RAPIDS Deployment: Custom RAPIDS Docker guide

dropna

A DataFrame method that removes rows or columns with nulls.

Objectives: 1.2

What NVIDIA says (1)

“any (default) drops rows (or columns) containing at least one null value. all drops only rows (or columns) containing all null values.”

— cuDF API: DataFrame.dropna

E

Edge list

A table with one row per edge, giving its source and destination nodes.

Objectives: 7.4

What NVIDIA says (1)

“This cudf.DataFrame contains columns storing edge source vertices, destination (or target following NetworkX’s terminology) vertices”

— cuGraph API: from_cudf_edgelist

Environment file

A text file, such as env.yaml or requirements.txt, that lists the packages a project needs.

Objectives: 8.1, 3.6

What NVIDIA says (2)

“To begin, you will need to create a few local files for your custom build: a Dockerfile and a configuration file ( env.yaml for conda or requirements.txt for pip).”

— RAPIDS Deployment: Custom RAPIDS Docker guide

“All dependencies (CUDA-X libraries, Prefect, MLflow, Triton client) are specified in environment.yml : $ conda env create -f environment.yml”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

ETL (Extract, transform, load)

The steps that pull raw data, clean and reshape it, and store it for analysis.

Objectives: 1.3, 5.4

What NVIDIA says (2)

“These tasks are often grouped under the term Extract, Transform, Load (ETL).”

— RAPIDS Accelerates Data Science End-to-End

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

Experiment tracking

Recording the settings, data and results of every training run so runs can be compared and repeated.

Objectives: 6.2

What NVIDIA says (2)

“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

“AI models require careful tracking through cycles of experiments, tuning and retraining.”

— What is MLOps?

Exploratory data analysis (EDA)

A first open-ended look at a dataset to learn its shape, gaps and patterns.

Objectives: 4.1

What NVIDIA says (1)

“Now, you understand the following data characteristics: Data types Dimensions of the dataset Number of sources garnering the dataset Dataset update frequency However, you must still explore whether this data has major gaps, either with missing or invalid data inputs.”

— Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

F

Feature

An input column a model learns from.

Objectives: 1.4, 3.2

What NVIDIA says (2)

“Next, numeric features must be normalized to prevent the model from being biased by the variable scales.”

— Accelerated Data Analytics: Machine Learning with GPU-Accelerated pandas and scikit-learn

“feature_importances_ ndarray of shape (n_features,) The impurity-based feature importances.”

— cuML API: RandomForestClassifier

Feature engineering

Creating or reshaping input columns so a model can learn better.

Objectives: 1.4, 3.2

What NVIDIA says (2)

“Next, numeric features must be normalized to prevent the model from being biased by the variable scales.”

— Accelerated Data Analytics: Machine Learning with GPU-Accelerated pandas and scikit-learn

“Then transform the ‘date’ column into an ‘hour’ feature, as weather patterns often correlate with the time of day.”

— Accelerated Data Analytics: Machine Learning with GPU-Accelerated pandas and scikit-learn

Feature importance

A score for how much each feature contributes to a model.

Objectives: 3.2

What NVIDIA says (1)

“feature_importances_ ndarray of shape (n_features,) The impurity-based feature importances.”

— cuML API: RandomForestClassifier

fillna

A DataFrame method that replaces nulls with chosen values.

Objectives: 1.2

What NVIDIA says (1)

“A dict can be used to provide different values to fill nulls in different columns.”

— cuDF API: DataFrame.fillna

Floating-point determinism

Getting bit-identical results on every run; parallel float sums may differ slightly, so compare with a tolerance.

Objectives: 3.6

What NVIDIA says (2)

“This impacts the determinism of floating-point operations because floating-point arithmetic is non-associative, that is, (a + b) + c is not necessarily equal to a + (b + c) .”

— cuDF: Comparison of cuDF and pandas

“If you need to compare floating point results, you should typically do so using the functions provided in the cudf.testing module, which allow you to compare values up to a desired precision.”

— cuDF: Comparison of cuDF and pandas

Forecast horizon

How many future steps a forecast predicts.

Objectives: 7.1

What NVIDIA says (1)

“In contrast, direct multi-step forecasting uses a separate model to predict each future value in your forecast horizon.”

— Accelerating Time Series Forecasting with RAPIDS cuML

G

Git

A version control system that records every change to a set of files.

Objectives: 8.5

What NVIDIA says (2)

“Key Concepts # Project A Git repository under management by AI Workbench.”

— NVIDIA AI Workbench: Introduction

“Both code and environment are Git managed with changes detected and surfaced in the Desktop App.”

— NVIDIA AI Workbench: Introduction

Git hosting service

A website, such as GitHub or GitLab, that stores repositories so teams can share them.

Objectives: 8.3, 8.5

What NVIDIA says (2)

“AI Workbench provides integration with GitHub for project collaboration and version control.”

— NVIDIA AI Workbench: Glossary

“AI Workbench supports both GitLab.com and self-hosted GitLab instances.”

— NVIDIA AI Workbench: Glossary

Git LFS (Large File Storage)

A Git extension that stores large files outside normal history and keeps pointers in the repository.

Objectives: 8.5

What NVIDIA says (2)

“A Git extension for versioning large files efficiently.”

— NVIDIA AI Workbench: Glossary

“AI Workbench automatically configures certain directories (like data/ and models/ ) to use Git LFS.”

— NVIDIA AI Workbench: Glossary

Graph

A set of nodes (entities) joined by edges (relationships).

Objectives: 7.4

What NVIDIA says (1)

“A graph consists of nodes or vertices (representing the entities in the system) that are connected by edges (representing relationships between those entities).”

— NVIDIA Glossary: Graph Analytics

Graph layout

Positions for drawing nodes so that the network structure is easy to see.

Objectives: 7.5

What NVIDIA says (1)

“ForceAtlas2 is a continuous graph layout algorithm for handy network visualization.”

— cuGraph API: force_atlas2

Graphics processing unit (GPU)

A processor with thousands of small cores that work on many pieces of a problem at once.

Objectives: 5.2, 5.3

What NVIDIA says (2)

“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”

— What's the Difference Between a CPU and a GPU?

“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”

— What Is Accelerated Computing?

Trying every combination of listed hyperparameter values.

Objectives: 2.4

What NVIDIA says (2)

“The grid search will take place over |n_estimators| x |max_depth| which is 3 x 3 = 9.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

“As you have probably guessed, the grid size grows rapidly as the number of parameters and their search space increases.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

Group-by aggregation

Splitting rows into groups by a key and computing a summary, such as a sum, for each group.

Objectives: 1.1

What NVIDIA says (1)

“With RAPIDS, all you have to do to use this code to run on a GPU and enjoy the interactive querying of data is to change the import statement.”

— Beginner's Guide to GPU-Accelerated DataFrames for pandas Users

H

Heat map

A grid colored by value, used to compare two categories at once.

Objectives: 4.3

What NVIDIA says (1)

“An hvPlot heat map showing trips by hour and day of week, per month Adding a widget for interactivity enables scrubbing through the months to search for patterns over a full year (Figure 2).”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Histogram

A chart of how often values fall into each range.

Objectives: 4.3

What NVIDIA says (1)

“An hvPlot histogram of trip durations generated with the Divvy dataset In this instance, the vast majority of bike trips appear under 20 minutes.”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Host and device

The CPU with its memory (host) and the GPU with its memory (device); data must be copied between them.

Objectives: 5.3

What NVIDIA says (2)

“Minimize the amount of data transferred between host and device when possible, even if that means running kernels on the GPU that get little or no speed-up compared to running them on the host CPU.”

— How to Optimize Data Transfers in CUDA C/C++

“The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”

— How to Optimize Data Transfers in CUDA C/C++

Hyperparameter

A setting chosen before training, such as tree depth.

Objectives: 2.4, 5.6

What NVIDIA says (2)

“Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

“This is particularly important because data scientists typically run the algorithm not just once, but many times in order to tune hyperparameters (such as learning rate or tree depth) and find the best accuracy.”

— Gradient Boosting, Decision Trees and XGBoost with CUDA

Hyperparameter optimization (HPO)

Searching for the hyperparameter values that give the best validation score.

Objectives: 2.4

What NVIDIA says (1)

“Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

I

Interpolation

Estimating a missing value from the known values around it.

Objectives: 7.2

What NVIDIA says (2)

“Parameters : method str, default ‘linear’ Interpolation technique to use.”

— cuDF API: DataFrame.interpolate

“‘index’, ‘values’: linearly interpolate using the index as an x-axis.”

— cuDF API: DataFrame.interpolate

J

Join (merge)

Combining two tables by matching values in a key column.

Objectives: 1.1, 3.4

What NVIDIA says (2)

“df_merged = df_a . merge ( df_b , on = [ 'key' ], how = 'left' )”

— cuDF API: DataFrame.merge

“Note that the dataframe order is not maintained , but may be restored post-merge by sorting by the index.”

— cuDF: 10 Minutes to cuDF and Dask cuDF

Jupyter notebook

An interactive document that mixes code, output and notes, run cell by cell.

Objectives: 5.1

What NVIDIA says (2)

“Just %load_ext cudf.pandas in Jupyter, or pass -m cudf.pandas on the command line.”

— cuDF: cudf.pandas

“Additionally, you can use Python magic commands like %%time and %%timeit to enable benchmarks of specific code blocks that facilitate direct comparisons of runtime between pandas (CPU) and the cuDF accelerator for pandas (GPU).”

— Get Started with GPU Acceleration for Data Science

K

k-fold cross-validation

Splitting data into k parts and validating on each part once while training on the rest.

Objectives: 2.5

What NVIDIA says (2)

“Each fold is then used once as a validation set while the k - 1 remaining folds form the training set.”

— cuML API: KFold

“Split dataset into k consecutive folds (without shuffling by default).”

— cuML API: KFold

KvikIO

A library cuDF uses for fast, parallel file input and output, including GPUDirect Storage.

Objectives: 1.7

What NVIDIA says (1)

“cuDF leverages the KvikIO library for high-performance I/O features, such as parallel I/O operations and NVIDIA Magnum IO GPUDirect Storage (GDS).”

— cuDF: Input / Output

L

Lag feature

A past value of the target used as an input, such as last week's sales.

Objectives: 7.1

What NVIDIA says (1)

“Lag features are useful because what happens in the past often influences what would happen in the future.”

— RAPIDS Deployment: Time series forecasting with HPO

LocalCUDACluster

A Dask-CUDA cluster on one machine with one worker per GPU.

Objectives: 3.6, 3.5

What NVIDIA says (2)

“For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”

— cuML: Multi-GPU with Dask

“This automatically creates one worker per available GPU.”

— cuML: Multi-GPU with Dask

M

Mean absolute error (MAE)

The average size of errors, ignoring sign; less swayed by outliers than MSE.

Objectives: 2.6

What NVIDIA says (1)

“Some of those might be outliers, so MSE is not robust to their presence.”

— A Comprehensive Overview of Regression Evaluation Metrics

MLflow

An open-source tool that records experiment settings, metrics and models, and keeps a model registry.

Objectives: 6.2, 6.5

What NVIDIA says (2)

“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

“The champion alias always points to the current production model: In a typical production cycle, start with a full pipeline run to establish a baseline champion.”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

MLOps (machine learning operations)

Practices for building, deploying and running machine learning reliably in production.

Objectives: 6.1, 6.2

What NVIDIA says (2)

“AI models require careful tracking through cycles of experiments, tuning and retraining.”

— What is MLOps?

“Storing infrastructure configurations in version control so production environments can be quickly replicated, whether for fault-tolerance or to create a replica environment for development.”

— What Is MLOps? (NVIDIA Glossary)

Model artifact

A file produced by a run, such as a saved model, that is stored with the run record.

Objectives: 6.5, 6.2

What NVIDIA says (1)

“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Model monitoring

Tracking a deployed model to catch slowdowns and drops in quality.

Objectives: 6.1, 6.4

What NVIDIA says (2)

“Monitoring a machine learning model after deployment is vital, as models can break and degrade in production.”

— A Guide to Monitoring Machine Learning Models in Production

“Examples include: Latency IO/memory/disk use System reliability (uptime) Auditability”

— A Guide to Monitoring Machine Learning Models in Production

Model registry

A store of model versions with names and aliases, such as the current production "champion".

Objectives: 6.5

What NVIDIA says (2)

“The champion alias always points to the current production model: In a typical production cycle, start with a full pipeline run to establish a baseline champion.”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Model serialization

Saving a trained model to a file so it can be loaded later, for example with pickle.

Objectives: 6.3

What NVIDIA says (2)

“Single GPU Model Serialization # All single-GPU cuML estimators support serialization using standard Python libraries.”

— cuML: Pickling cuML models for persistence

“Security Warning # Only unpickle or deserialize models from trusted sources.”

— cuML: Pickling cuML models for persistence

MSE and RMSE (mean squared error, root mean squared error)

The average squared error, and its square root in the units of the target.

Objectives: 2.6

What NVIDIA says (2)

“Some of those might be outliers, so MSE is not robust to their presence.”

— A Comprehensive Overview of Regression Evaluation Metrics

“Other than the scale, RMSE has the same properties as MSE.”

— A Comprehensive Overview of Regression Evaluation Metrics

Multi-Instance GPU (MIG)

A feature that splits one GPU into several isolated instances that each act like a separate GPU.

Objectives: 6.6

What NVIDIA says (1)

“Multi-Instance GPU is a technology that allows partitioning a single GPU into multiple instances, making each one seem as a completely independent GPU.”

— RAPIDS Deployment: Multi-Instance GPU (MIG)

N

Null (missing value) (NA)

An empty entry where a value is unknown or absent.

Objectives: 1.2, 4.1

What NVIDIA says (2)

“cudf supports having missing values in all dtypes.”

— cuDF: Working with missing data

“To detect missing values, you can use isna() and notna() functions.”

— cuDF: Working with missing data

NumPy

A Python library for fast multi-dimensional arrays and math on whole arrays.

Objectives: 5.1

What NVIDIA says (2)

“What is NumPy NumPy is a powerful, well-optimized, free open-source library for the Python programming language, adding support for large, multi-dimensional arrays (also called matrices or tensors).”

— NVIDIA Glossary: NumPy

“As the core library for scientific computing, NumPy is the base for libraries such as Pandas, Scikit-learn , and SciPy .”

— NVIDIA Glossary: NumPy

NVIDIA AI Workbench

An NVIDIA tool that manages data science projects in containers with Git-managed code and environments.

Objectives: 8.1, 8.5

What NVIDIA says (2)

“AI Workbench provides reproducibility by managing software, containers and Git repositories.”

— NVIDIA AI Workbench: Introduction

“Versioned configuration files let you define and adapt environments to machines and users.”

— NVIDIA AI Workbench: Introduction

nvidia-smi (NVIDIA System Management Interface)

A command-line tool that lists NVIDIA GPUs and reports their status and use.

Objectives: 8.4

What NVIDIA says (1)

“nvidia-smi (also NVSMI) provides monitoring and management capabilities for each of NVIDIA's Tesla, Quadro, GRID and GeForce devices”

— nvidia-smi documentation

O

One-hot encoding

Turning a category column into one 0/1 column per category.

Objectives: 1.4

What NVIDIA says (2)

“cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”

— cuDF API: get_dummies

“However, dropping one category breaks the symmetry of the original representation and can therefore induce a bias in downstream models, for instance for penalized linear classification or regression models.”

— cuML API: OneHotEncoder

ONNX (Open Neural Network Exchange)

A portable model file format that runs without the original training library.

Objectives: 6.3

What NVIDIA says (2)

“Use as_sklearn() to convert the cuML model to a scikit-learn estimator, then pass it to skl2onnx.convert_sklearn() .”

— cuML: Pickling cuML models for persistence

“The resulting .onnx file can be loaded with ONNX Runtime for inference on both CPU and GPU, with no cuML dependency at inference time.”

— cuML: Pickling cuML models for persistence

Optuna

A framework that automates hyperparameter search and can run trials in parallel.

Objectives: 2.4

What NVIDIA says (2)

“Optuna is a lightweight framework for automatic hyperparameter optimization.”

— RAPIDS Deployment: Optuna HPO with RAPIDS

“By simply wrapping the objective function with Optuna, we can perform a parallel-distributed HPO search over a search space as we’ll see in this notebook.”

— RAPIDS Deployment: Optuna HPO with RAPIDS

Overfitting

When a model learns noise in the training data and does poorly on new data.

Objectives: 3.3, 5.6

What NVIDIA says (2)

“Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”

— Gradient Boosting, Decision Trees and XGBoost with CUDA

“The presence of a large number of trees also reduces the problem of overfitting, which occurs when a model incorporates too much “noise” in the training data and makes poor decisions as a result.”

— NVIDIA Glossary: Random Forest

P

p-value

The chance of seeing a result this extreme if there were really no effect; small values suggest a real effect.

Objectives: 4.4

What NVIDIA says (1)

“The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”

— Improving Player Performance with Low Latency as Evident from FPS Aim Trainer Experiments

PageRank

An importance score for each node, higher when important nodes link to it.

Objectives: 7.5

What NVIDIA says (1)

“Find the PageRank score for every vertex in a graph.”

— cuGraph API: pagerank

pandas

A Python library for working with tables of data, called DataFrames.

Objectives: 5.1, 1.1

What NVIDIA says (2)

“A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable, and each row contains a set of values corresponding to those variables.”

— What Is Pandas and Why Does it Matter?

“With its support for structured data formats like tables, matrices, and time series, the pandas Python API provides tools to process messy or raw datasets into clean, structured formats ready for analysis.”

— What Is Pandas and Why Does it Matter?

Parallel processing

Splitting a job into many pieces that run at the same time.

Objectives: 5.2, 5.5

What NVIDIA says (2)

“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”

— What's the Difference Between a CPU and a GPU?

“Dask is an open-source library designed to provide parallelism to the existing Python stack.”

— NVIDIA Glossary: Dask

Parquet

A columnar file format that stores data by column, so tools can read only the columns they need.

Objectives: 1.7

What NVIDIA says (2)

“columns list, default None If not None, only these columns will be read.”

— cuDF API: read_parquet

“If not None, specifies a filter predicate used to filter out row groups using statistics stored for each row group as Parquet metadata.”

— cuDF API: read_parquet

Partitioned dataset

A dataset split into folders by column values, such as one folder per year.

Objectives: 1.7

What NVIDIA says (1)

“partition_cols list, optional, default None Column names by which to partition the dataset Columns are partitioned in the order they are given”

— cuDF API: DataFrame.to_parquet

Personal identifiable information (PII)

Data that can identify a person, such as a name or phone number.

Objectives: 1.2

What NVIDIA says (1)

“Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”

— Enhancing Generative AI Model Accuracy with NVIDIA NeMo Curator

Pinned memory (page-locked memory)

Host memory the operating system cannot move, so the GPU can copy from it directly and faster.

Objectives: 5.3

What NVIDIA says (2)

“Higher bandwidth is possible between the host and the device when using page-locked (or “pinned”) memory.”

— How to Optimize Data Transfers in CUDA C/C++

“The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”

— How to Optimize Data Transfers in CUDA C/C++

pip

Python's standard package installer; it cannot install system-level CUDA libraries.

Objectives: 8.2

What NVIDIA says (1)

“Since pip cannot install system-level CUDA libraries, CuPy expects these libraries to already be present in the system environment.”

— RAPIDS Deployment: Custom RAPIDS Docker guide

Plotly Dash

A Python framework for building interactive web dashboards.

Objectives: 4.2

What NVIDIA says (2)

“Plotly Dash enables data scientists to recast complex data and machine learning workflows as more accessible web applications.”

— Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

“The use of Plotly’s Dash, RAPIDS, and Data shader allows users to build viz dashboards that both render datasets of 300 million+ rows and remain highly interactive without the need for precomputed aggregations.”

— Making a Plotly Dash Census Viz Powered by RAPIDS

Polars GPU engine

A cuDF-powered engine that runs Polars lazy queries on the GPU and falls back to the CPU when needed.

Objectives: 1.7

What NVIDIA says (2)

“cuDF provides GPU-accelerated execution engines for Python users of the Polars Lazy API.”

— cuDF: Polars GPU engine

“If it is not, the execution transparently falls back to the standard Polars engine and runs on the CPU.”

— cuDF: Polars GPU engine

Precision

Of the cases the model flagged as positive, the share that really are.

Objectives: 2.6

What NVIDIA says (1)

“The precision is the ratio tp / (tp + fp) where tp is the number of true positives and fp the number of false positives.”

— cuML API: precision_recall_curve

Prediction drift

A change in the distribution of a model's outputs, watched when true labels are not available.

Objectives: 6.4

What NVIDIA says (2)

“Prediction drift: When it is not possible to acquire ground truth labels, predictions must be monitored.”

— A Guide to Monitoring Machine Learning Models in Production

“If there is a drastic change in the distribution of predictions, something has potentially gone wrong.”

— A Guide to Monitoring Machine Learning Models in Production

Prefect

A workflow orchestrator that chains pipeline stages, tracks runs, schedules them and handles failures.

Objectives: 3.5, 6.1

What NVIDIA says (2)

“Prefect orchestrates the pipeline stages, tracks runs, and enables scheduling”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

“Orchestration: Prefect flows chain the stages together, track runs, and handle failures”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Principal component analysis (PCA)

A method that combines features into a few linear components that keep the most variance.

Objectives: 1.6

What NVIDIA says (1)

“PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”

— cuML API: PCA

R

R-squared (R²)

The share of the variation in the target that a regression model explains.

Objectives: 2.3

What NVIDIA says (2)

“R-squared (R²), also known as the coefficient of determination, represents the proportion of variance explained by a model.”

— A Comprehensive Overview of Regression Evaluation Metrics

“Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”

— cuML API: r2_score

Random forest

Many decision trees trained on random samples whose votes are combined.

Objectives: 2.2, 3.3

What NVIDIA says (2)

“Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”

— NVIDIA Glossary: Random Forest

“Whereas decision trees are based upon a fixed set of features, and often overfit, randomness is critical to the success of the forest.”

— NVIDIA Glossary: Random Forest

Trying randomly sampled hyperparameter combinations instead of all of them.

Objectives: 2.4

What NVIDIA says (1)

“Random Search replaces the exhaustive nature of the search from before with a random selection of parameters over the specified space.”

— RAPIDS Deployment: XGBoost and Random Forest GPU HPO with Dask

RAPIDS

NVIDIA's open-source GPU libraries for data science, covering loading, ETL, training and analysis.

Objectives: 5.4, 3.1

What NVIDIA says (2)

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”

— RAPIDS Accelerates Data Science End-to-End

RAPIDS Accelerator for Apache Spark

A plugin that runs Apache Spark jobs on GPUs without code changes.

Objectives: 1.3

What NVIDIA says (2)

“Run your existing Apache Spark applications with no code change.”

— cuDF for Apache Spark: Overview

“The optimization came from using the RAPIDS accelerator for Apache Spark, an open-source library that enables GPU-accelerated ETL and feature engineering.”

— Scaling Data Pipelines: AT&T Optimizes Speed, Cost, and Efficiency with GPUs

Recall

Of the real positive cases, the share the model found.

Objectives: 2.6

What NVIDIA says (2)

“The recall is the ratio tp / (tp + fn) where tp is the number of true positives and fn the number of false negatives.”

— cuML API: precision_recall_curve

“The recall is intuitively the ability of the classifier to find all the positive samples.”

— cuML API: precision_recall_curve

Regression

Predicting a number, such as a price.

Objectives: 2.2

What NVIDIA says (1)

“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”

— NVIDIA Glossary: Linear Regression and Logistic Regression

Regularization

A penalty on model complexity that helps prevent overfitting.

Objectives: 3.3

What NVIDIA says (2)

“Ridge extends LinearRegression by providing L2 regularization on the coefficients when predicting response y with a linear combination of the predictors in X.”

— cuML API: Ridge

“Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”

— Gradient Boosting, Decision Trees and XGBoost with CUDA

Repository

A project folder plus its full Git history.

Objectives: 8.5

What NVIDIA says (1)

“Key Concepts # Project A Git repository under management by AI Workbench.”

— NVIDIA AI Workbench: Introduction

Resampling

Converting a time series to a new regular frequency, such as one value per minute.

Objectives: 7.2

What NVIDIA says (2)

“Parameters : rule: str The offset string representing the frequency to use.”

— cuDF API: DataFrame.resample

“First, we create a time series with 1 minute intervals”

— cuDF API: DataFrame.resample

Rolling mean

The average over a sliding window of recent rows, used to smooth noise.

Objectives: 4.5, 7.1

What NVIDIA says (2)

“Parameters : window int, offset or a BaseIndexer subclass Size of the window, i.e., the number of observations used to calculate the statistic.”

— cuDF API: DataFrame.rolling

“Rolling window statistics are statistics (e.g. mean, standard deviation) over a time duration in the past.”

— RAPIDS Deployment: Time series forecasting with HPO

S

Sampling

Taking a random subset of rows, for example for quick exploration.

Objectives: 1.6

What NVIDIA says (2)

“This function will always produce the same sample given an identical random_state .”

— cuDF API: DataFrame.sample

“frac float, optional Fraction of axis items to return.”

— cuDF API: DataFrame.sample

scikit-learn

A popular CPU machine learning library for Python with a standard fit and predict API.

Objectives: 2.1

What NVIDIA says (1)

“cuML estimators look and feel just like scikit-learn estimators .”

— cuML: Introduction

SHAP (SHapley Additive exPlanations)

A method that explains a single prediction by how much each feature pushed it.

Objectives: 3.2

What NVIDIA says (1)

“cuML’s SHAP based explainers accelerate the algorithmic part of SHAP.”

— cuML API: KernelExplainer

Standard scaling

Rescaling a numeric feature to mean 0 and standard deviation 1.

Objectives: 1.4, 3.2

What NVIDIA says (2)

“Standardize features by removing the mean and scaling to unit variance”

— cuML API: StandardScaler

“Mean and standard deviation are then stored to be used on later data using transform() .”

— cuML API: StandardScaler

Stratified split

A train-test split that keeps the same class ratio in both parts.

Objectives: 1.5

What NVIDIA says (1)

“If not None, data is split in a stratified fashion, using this as the class labels.”

— cuML API: train_test_split

Supervised learning

Learning from examples that come with the correct answer (labels).

Objectives: 2.2

What NVIDIA says (2)

“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”

— NVIDIA Glossary: Linear Regression and Logistic Regression

“Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”

— NVIDIA Glossary: Linear Regression and Logistic Regression

Synthetic data

Generated data that mimics real data, used to fill gaps or protect privacy.

Objectives: 1.5, 3.4

What NVIDIA says (2)

“Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”

— NVIDIA Glossary: Synthetic Data Generation

“Synthetically generated data can augment existing data for larger, more representative datasets.”

— NVIDIA Glossary: Synthetic Data Generation

T

Target encoding

Replacing each category with the average target value for that category.

Objectives: 1.4

What NVIDIA says (2)

“The input data is grouped by the columns Xs and the aggregated mean value of Y of each group is calculated to replace each value of Xs .”

— cuML API: TargetEncoder

“Several optimizations are applied to prevent label leakage and parallelize the execution.”

— cuML API: TargetEncoder

Test set

Data held back from training and used once to estimate performance on new data.

Objectives: 2.3

What NVIDIA says (1)

“test_size float or int, default=None If float, should be between 0.0 and 1.0 and represent the proportion of the dataset to include”

— cuML API: train_test_split

Time series

Data points recorded in time order.

Objectives: 7.1

What NVIDIA says (2)

“Great care must be taken when defining cross-validation folds for time-series data.”

— RAPIDS Deployment: Time series forecasting with HPO

“First, we create a time series with 1 minute intervals”

— cuDF API: DataFrame.resample

U

UMAP (Uniform Manifold Approximation and Projection)

A non-linear method that maps high-dimensional data to two or three dimensions for plotting.

Objectives: 1.6

What NVIDIA says (2)

“UMAP is a popular dimension reduction algorithm used in fields like bioinformatics, NLP topic modeling, and ML preprocessing.”

— Even Faster and More Scalable UMAP on the GPU with RAPIDS cuML

“It works by creating a k-nearest neighbors (k-NN) graph, which is known in literature as an all-neighbors graph, to build a fuzzy topological representation of the data, which is used to embed high-dimensional data into lower dimensions.”

— Even Faster and More Scalable UMAP on the GPU with RAPIDS cuML

Underfitting

When a model is too simple to capture the pattern, even on training data.

Objectives: 5.6, 3.3

What NVIDIA says (2)

“Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”

— What Is XGBoost and Why Does It Matter?

“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”

— NVIDIA Glossary: Random Forest

Unsupervised learning

Finding patterns in data that has no labels.

Objectives: 2.2

What NVIDIA says (2)

“Unsupervised learning algorithms attempt to ‘learn’ patterns in unlabeled data sets, discovering similarities, or regularities.”

— NVIDIA Glossary: K-Means

“Common unsupervised tasks include clustering and association.”

— NVIDIA Glossary: K-Means

W

Weights & Biases (W&B)

A platform for tracking, debugging and reproducing ML experiments.

Objectives: 6.2

What NVIDIA says (1)

“Weights & Biases : W&B’s tools help many ML users build better models faster, debugging and reproducing their models with just a few lines of code”

— What is MLOps?

X

XGBoost (Extreme Gradient Boosting)

A scalable library for gradient-boosted decision trees that can train on GPUs.

Objectives: 2.1

What NVIDIA says (1)

“XGBoost , which stands for Extreme Gradient Boosting, is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”

— What Is XGBoost and Why Does It Matter?