NCA-ADS glossary
The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.
A
- Accelerated computing
Using specialized hardware such as GPUs to speed up work through parallel processing.
What NVIDIA says (1)
“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”
- Accuracy
The share of predictions that are correct.
What NVIDIA says (2)
“Accuracy score is the ratio of correct predictions to the total number of predictions.”
“Accuracy classification score.”
B
- Bagging
Training many models on random samples and averaging them to reduce variance.
What NVIDIA says (2)
“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”
“Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”
- Benchmarking
Timing the same workload under different setups, such as CPU and GPU, to compare them fairly.
What NVIDIA says (2)
“In the notebook demo below, we compare benchmarking results to show how GPU can accelerate HPO tuning jobs relative to CPU.”
“Additionally, you can use Python magic commands like %%time and %%timeit to enable benchmarks of specific code blocks that facilitate direct comparisons of runtime between pandas (CPU) and the cuDF accelerator for pandas (GPU).”
- Binning
Grouping a continuous number into ranges, such as age bands.
What NVIDIA says (1)
“Bin continuous data into intervals.”
- Boosting
Training models one after another, each fixing the errors of the last, to reduce bias.
What NVIDIA says (1)
“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”
C
- Central processing unit
A computer's general-purpose processor, with a few fast cores that suit step-by-step work and small data.
What NVIDIA says (2)
“Small data sizes may be slower on GPU than CPU, because of the cost of data transfers.”
“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”
- Class imbalance
When one class is much rarer than the others in the data.
What NVIDIA says (2)
“The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”
“Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”
- Classification
Predicting a category, such as spam or not spam.
What NVIDIA says (1)
“Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”
- Clustering
Grouping similar unlabeled rows together.
What NVIDIA says (2)
“Common unsupervised tasks include clustering and association.”
“DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”
- conda
A package and environment manager that installs Python and non-Python packages, including CUDA libraries.
What NVIDIA says (2)
“When installing libraries via conda, the package manager automatically pulls the required CUDA runtime libraries alongside CuPy and other dependencies, providing complete dependency management in a single installation step.”
“The defaults channel is not supported by these packages, which are built to be compatible with dependencies from the conda-forge channel.”
- Conda channel
A source that conda downloads packages from, such as conda-forge.
What NVIDIA says (1)
“The defaults channel is not supported by these packages, which are built to be compatible with dependencies from the conda-forge channel.”
- Confidence interval
A range that likely contains the true value; more samples make it narrower.
What NVIDIA says (2)
“A single success rate on N rollouts tells you almost nothing about how confident you should be in a policy’s true performance.”
“Narrowing the confidence interval from 10 to 2 percentage points requires roughly 15x more rollouts (70 to 1,030).”
- Confusion matrix
A table that counts each true class against each predicted class.
What NVIDIA says (2)
“Compute confusion matrix to evaluate the accuracy of a classification.”
“Normalizes confusion matrix over the true (rows), predicted (columns) conditions or al”
- Container
A packaged environment with an app and its dependencies that runs the same on any host.
What NVIDIA says (2)
“These size reductions result in faster container pulls and deployments, reduced storage costs in container registries, lower bandwidth usage in distributed environments, and quicker startup times for containerized applications.”
“AI Workbench provides reproducibility by managing software, containers and Git repositories.”
- Correlation
A measure of how two variables move together; it does not prove cause.
What NVIDIA says (2)
“pearson : Standard correlation coefficient spearman : Spearman rank correlation”
“However, it’s important to remember that correlation and causation are two different things.”
- CPU fallback
Running an operation on the CPU when the GPU path does not support it, at the cost of extra data copies.
What NVIDIA says (2)
“If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”
“CPU fallback preserves compatibility, but frequent transitions between CPU and GPU execution may reduce the overall speedup.”
- Cross-filtering
Linked charts where selecting data in one filters all the others.
What NVIDIA says (1)
“Instead of creating several individual group by and query operations, a cuxfilter dashboard can simply cross-link numerous charts to quickly find patterns or anomalies (Figure 5).”
- Cross-validation
Rotating which part of the data is held out, to get a steadier performance estimate.
What NVIDIA says (2)
“Cross-validation is often used to more accurately estimate the performance of the models in the search process.”
“Cross-validation is the method of splitting the training set into complementary subsets and performing training on one of the subsets, then predicting the models performance on the other.”
- CUDA driver
The NVIDIA system software that lets programs use the GPU; it must support the CUDA version your packages need.
What NVIDIA says (2)
“You will have to ensure the CUDA driver on your machine supports the CUDA version you are trying to install with conda.”
“If conda has incorrectly identified the CUDA driver, you can override by setting the CONDA_OVERRIDE_CUDA environment variable.”
- CUDA_VISIBLE_DEVICES
An environment variable that sets which GPUs a program can see.
What NVIDIA says (1)
“To specify a device to run on, we recommend using the CUDA_VISIBLE_DEVICES ( doc ) environment variable.”
- cuDF
A GPU DataFrame library with a pandas-style API for loading, joining, aggregating and filtering data.
What NVIDIA says (1)
“cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”
- cudf.pandas
A cuDF mode that runs existing pandas code on the GPU without code changes and falls back to the CPU when needed.
What NVIDIA says (2)
“It supports 100% of the Pandas API , using the GPU for supported operations, and automatically falling back to pandas for other operations.”
“If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”
- cuGraph
A RAPIDS library for graph analytics on the GPU.
What NVIDIA says (2)
“A GPU Graph Object (Base class of other graph types)”
“Find the PageRank score for every vertex in a graph.”
- cuML
A GPU machine learning library whose estimators work like scikit-learn estimators.
What NVIDIA says (2)
“cuML estimators look and feel just like scikit-learn estimators .”
“NVIDIA cuML is a GPU-accelerated machine learning library for Python with a scikit-learn compatible API.”
- cuml.accel
A cuML mode that runs existing scikit-learn code on the GPU without code changes.
What NVIDIA says (2)
“When running a script, use the cuml.accel command-line interface: python -m cuml.accel script.py In Jupyter or IPython, load the extension before other imports: % load_ext cuml.accel”
“Enable cuml.accel before importing scikit-learn, UMAP, or HDBSCAN.”
- cuxfilter
A RAPIDS library for GPU dashboards whose charts filter each other.
What NVIDIA says (1)
“Instead of creating several individual group by and query operations, a cuxfilter dashboard can simply cross-link numerous charts to quickly find patterns or anomalies (Figure 5).”
D
- Dask
A Python library that runs familiar tools such as pandas in parallel across cores, GPUs and machines.
What NVIDIA says (2)
“Dask is an open-source library designed to provide parallelism to the existing Python stack.”
“However, unlike Apache Spark, it does not introduce a new API but provides a familiar programming interface of tools found in the PyData ecosystem, like pandas, scikit-learn, NetworkX, and more.”
- Dask cuDF
Dask DataFrames backed by cuDF, for data split into partitions across GPUs.
What NVIDIA says (2)
“When installed, Dask cuDF is automatically registered as the "cudf" dataframe backend for Dask DataFrame .”
“You must also deploy a dask.distributed cluster to leverage multiple GPUs.”
- Data augmentation
Adding new rows or derived columns to make training data richer.
What NVIDIA says (2)
“Synthetically generated data can augment existing data for larger, more representative datasets.”
“Augmenting the data using RAPIDS cuSpatial to quickly calculate distances also shows that most trips are relatively short.”
- Data drift
A change in the input data between training and production.
What NVIDIA says (1)
“Data drift : Changes in distribution between the training data and production data can be monitored to check for drift: this is done by detecting changes in the statistical properties of feature values over time.”
- Data flywheel
A loop where production feedback and new data keep improving a deployed model.
What NVIDIA says (2)
“AI data flywheels work by creating a loop where AI models continuously improve by learning from the latest institutional knowledge and user feedback.”
“As the system interacts with the environment, it collects feedback and new data, which are then used to refine and enhance the backbone models powering the AI workflows.”
- Data governance
Rules for using data lawfully and responsibly, including privacy and consent.
What NVIDIA says (2)
“In addition to complying with privacy and consumer protection laws, trustworthy AI models are tested for safety, security and mitigation of unwanted bias.”
“Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”
- Data science pipeline
The chain of steps from loading data to training and serving a model.
What NVIDIA says (2)
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
“Each stage can also run independently for ad-hoc experiments.”
- DataFrame
A two-dimensional table where each column is one variable and each row is one record.
What NVIDIA says (1)
“A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable, and each row contains a set of values corresponding to those variables.”
- Datashader
A library that turns millions of points into an accurate image by aggregating them first.
What NVIDIA says (2)
“The Datashader library directly supports cuDF and can rapidly render over millions of aggregated points.”
“Datapoint rendering displaying high-resolution patterns is precisely what Datashader is designed for.”
- DBSCAN
A clustering method that finds dense regions of any shape and marks sparse points as noise.
What NVIDIA says (2)
“DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”
“This also allows DBSCAN to be robust to noise.”
- Deduplication
Removing repeated records so they do not skew training.
What NVIDIA says (2)
“Training models on these datasets without proper processing can result in higher training time and lower model quality.”
“Determines which duplicates (if any) to keep. - ‘first’ : Drop duplicates except for the first occurrence.”
- Descriptive statistics
Summary numbers such as count, mean, standard deviation and percentiles.
What NVIDIA says (2)
“For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”
“The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”
- Dimensionality reduction
Representing data with fewer columns while keeping most of its structure.
What NVIDIA says (2)
“PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”
“Reduce dimensionality through Gaussian random projection.”
- Dockerfile
A text recipe for building a container image.
What NVIDIA says (1)
“To begin, you will need to create a few local files for your custom build: a Dockerfile and a configuration file ( env.yaml for conda or requirements.txt for pip).”
- dropna
A DataFrame method that removes rows or columns with nulls.
What NVIDIA says (1)
“any (default) drops rows (or columns) containing at least one null value. all drops only rows (or columns) containing all null values.”
E
- Edge list
A table with one row per edge, giving its source and destination nodes.
What NVIDIA says (1)
“This cudf.DataFrame contains columns storing edge source vertices, destination (or target following NetworkX’s terminology) vertices”
- Environment file
A text file, such as env.yaml or requirements.txt, that lists the packages a project needs.
What NVIDIA says (2)
“To begin, you will need to create a few local files for your custom build: a Dockerfile and a configuration file ( env.yaml for conda or requirements.txt for pip).”
“All dependencies (CUDA-X libraries, Prefect, MLflow, Triton client) are specified in environment.yml : $ conda env create -f environment.yml”
- ETL
The steps that pull raw data, clean and reshape it, and store it for analysis.
What NVIDIA says (2)
“These tasks are often grouped under the term Extract, Transform, Load (ETL).”
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
- Experiment tracking
Recording the settings, data and results of every training run so runs can be compared and repeated.
What NVIDIA says (2)
“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”
“AI models require careful tracking through cycles of experiments, tuning and retraining.”
- Exploratory data analysis
A first open-ended look at a dataset to learn its shape, gaps and patterns.
What NVIDIA says (1)
“Now, you understand the following data characteristics: Data types Dimensions of the dataset Number of sources garnering the dataset Dataset update frequency However, you must still explore whether this data has major gaps, either with missing or invalid data inputs.”
F
- Feature
An input column a model learns from.
What NVIDIA says (2)
“Next, numeric features must be normalized to prevent the model from being biased by the variable scales.”
“feature_importances_ ndarray of shape (n_features,) The impurity-based feature importances.”
- Feature engineering
Creating or reshaping input columns so a model can learn better.
What NVIDIA says (2)
“Next, numeric features must be normalized to prevent the model from being biased by the variable scales.”
“Then transform the ‘date’ column into an ‘hour’ feature, as weather patterns often correlate with the time of day.”
- Feature importance
A score for how much each feature contributes to a model.
What NVIDIA says (1)
“feature_importances_ ndarray of shape (n_features,) The impurity-based feature importances.”
- fillna
A DataFrame method that replaces nulls with chosen values.
What NVIDIA says (1)
“A dict can be used to provide different values to fill nulls in different columns.”
- Floating-point determinism
Getting bit-identical results on every run; parallel float sums may differ slightly, so compare with a tolerance.
What NVIDIA says (2)
“This impacts the determinism of floating-point operations because floating-point arithmetic is non-associative, that is, (a + b) + c is not necessarily equal to a + (b + c) .”
“If you need to compare floating point results, you should typically do so using the functions provided in the cudf.testing module, which allow you to compare values up to a desired precision.”
- Forecast horizon
How many future steps a forecast predicts.
What NVIDIA says (1)
“In contrast, direct multi-step forecasting uses a separate model to predict each future value in your forecast horizon.”
G
- Git
A version control system that records every change to a set of files.
What NVIDIA says (2)
“Key Concepts # Project A Git repository under management by AI Workbench.”
“Both code and environment are Git managed with changes detected and surfaced in the Desktop App.”
- Git hosting service
A website, such as GitHub or GitLab, that stores repositories so teams can share them.
What NVIDIA says (2)
“AI Workbench provides integration with GitHub for project collaboration and version control.”
“AI Workbench supports both GitLab.com and self-hosted GitLab instances.”
- Git LFS
A Git extension that stores large files outside normal history and keeps pointers in the repository.
What NVIDIA says (2)
“A Git extension for versioning large files efficiently.”
“AI Workbench automatically configures certain directories (like data/ and models/ ) to use Git LFS.”
- Graph
A set of nodes (entities) joined by edges (relationships).
What NVIDIA says (1)
“A graph consists of nodes or vertices (representing the entities in the system) that are connected by edges (representing relationships between those entities).”
- Graph layout
Positions for drawing nodes so that the network structure is easy to see.
What NVIDIA says (1)
“ForceAtlas2 is a continuous graph layout algorithm for handy network visualization.”
- Graphics processing unit
A processor with thousands of small cores that work on many pieces of a problem at once.
What NVIDIA says (2)
“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”
“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”
- Grid search
Trying every combination of listed hyperparameter values.
What NVIDIA says (2)
“The grid search will take place over |n_estimators| x |max_depth| which is 3 x 3 = 9.”
“As you have probably guessed, the grid size grows rapidly as the number of parameters and their search space increases.”
- Group-by aggregation
Splitting rows into groups by a key and computing a summary, such as a sum, for each group.
What NVIDIA says (1)
“With RAPIDS, all you have to do to use this code to run on a GPU and enjoy the interactive querying of data is to change the import statement.”
H
- Heat map
A grid colored by value, used to compare two categories at once.
What NVIDIA says (1)
“An hvPlot heat map showing trips by hour and day of week, per month Adding a widget for interactivity enables scrubbing through the months to search for patterns over a full year (Figure 2).”
- Histogram
A chart of how often values fall into each range.
What NVIDIA says (1)
“An hvPlot histogram of trip durations generated with the Divvy dataset In this instance, the vast majority of bike trips appear under 20 minutes.”
- Host and device
The CPU with its memory (host) and the GPU with its memory (device); data must be copied between them.
What NVIDIA says (2)
“Minimize the amount of data transferred between host and device when possible, even if that means running kernels on the GPU that get little or no speed-up compared to running them on the host CPU.”
“The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”
- Hyperparameter
A setting chosen before training, such as tree depth.
What NVIDIA says (2)
“Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”
“This is particularly important because data scientists typically run the algorithm not just once, but many times in order to tune hyperparameters (such as learning rate or tree depth) and find the best accuracy.”
- Hyperparameter optimization
Searching for the hyperparameter values that give the best validation score.
What NVIDIA says (1)
“Hyperparameter optimization is the task of picking hyperparameters values of the model that provide the optimal results for the problem, as measured on a specific test dataset.”
I
- Interpolation
Estimating a missing value from the known values around it.
What NVIDIA says (2)
“Parameters : method str, default ‘linear’ Interpolation technique to use.”
“‘index’, ‘values’: linearly interpolate using the index as an x-axis.”
J
- Join (merge)
Combining two tables by matching values in a key column.
What NVIDIA says (2)
“df_merged = df_a . merge ( df_b , on = [ 'key' ], how = 'left' )”
“Note that the dataframe order is not maintained , but may be restored post-merge by sorting by the index.”
- Jupyter notebook
An interactive document that mixes code, output and notes, run cell by cell.
What NVIDIA says (2)
“Just %load_ext cudf.pandas in Jupyter, or pass -m cudf.pandas on the command line.”
“Additionally, you can use Python magic commands like %%time and %%timeit to enable benchmarks of specific code blocks that facilitate direct comparisons of runtime between pandas (CPU) and the cuDF accelerator for pandas (GPU).”
K
- k-fold cross-validation
Splitting data into k parts and validating on each part once while training on the rest.
What NVIDIA says (2)
“Each fold is then used once as a validation set while the k - 1 remaining folds form the training set.”
“Split dataset into k consecutive folds (without shuffling by default).”
- KvikIO
A library cuDF uses for fast, parallel file input and output, including GPUDirect Storage.
What NVIDIA says (1)
“cuDF leverages the KvikIO library for high-performance I/O features, such as parallel I/O operations and NVIDIA Magnum IO GPUDirect Storage (GDS).”
L
- Lag feature
A past value of the target used as an input, such as last week's sales.
What NVIDIA says (1)
“Lag features are useful because what happens in the past often influences what would happen in the future.”
- LocalCUDACluster
A Dask-CUDA cluster on one machine with one worker per GPU.
What NVIDIA says (2)
“For single-node, multi-GPU execution, use LocalCUDACluster from dask-cuda .”
“This automatically creates one worker per available GPU.”
M
- Mean absolute error
The average size of errors, ignoring sign; less swayed by outliers than MSE.
What NVIDIA says (1)
“Some of those might be outliers, so MSE is not robust to their presence.”
- MLflow
An open-source tool that records experiment settings, metrics and models, and keeps a model registry.
What NVIDIA says (2)
“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”
“The champion alias always points to the current production model: In a typical production cycle, start with a full pipeline run to establish a baseline champion.”
- MLOps
Practices for building, deploying and running machine learning reliably in production.
What NVIDIA says (2)
“AI models require careful tracking through cycles of experiments, tuning and retraining.”
“Storing infrastructure configurations in version control so production environments can be quickly replicated, whether for fault-tolerance or to create a replica environment for development.”
- Model artifact
A file produced by a run, such as a saved model, that is stored with the run record.
What NVIDIA says (1)
“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”
- Model monitoring
Tracking a deployed model to catch slowdowns and drops in quality.
What NVIDIA says (2)
“Monitoring a machine learning model after deployment is vital, as models can break and degrade in production.”
“Examples include: Latency IO/memory/disk use System reliability (uptime) Auditability”
- Model registry
A store of model versions with names and aliases, such as the current production "champion".
What NVIDIA says (2)
“The champion alias always points to the current production model: In a typical production cycle, start with a full pipeline run to establish a baseline champion.”
“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”
- Model serialization
Saving a trained model to a file so it can be loaded later, for example with pickle.
What NVIDIA says (2)
“Single GPU Model Serialization # All single-GPU cuML estimators support serialization using standard Python libraries.”
“Security Warning # Only unpickle or deserialize models from trusted sources.”
- MSE and RMSE
The average squared error, and its square root in the units of the target.
What NVIDIA says (2)
“Some of those might be outliers, so MSE is not robust to their presence.”
“Other than the scale, RMSE has the same properties as MSE.”
- Multi-Instance GPU
A feature that splits one GPU into several isolated instances that each act like a separate GPU.
What NVIDIA says (1)
“Multi-Instance GPU is a technology that allows partitioning a single GPU into multiple instances, making each one seem as a completely independent GPU.”
N
- Null (missing value)
An empty entry where a value is unknown or absent.
What NVIDIA says (2)
“cudf supports having missing values in all dtypes.”
“To detect missing values, you can use isna() and notna() functions.”
- NumPy
A Python library for fast multi-dimensional arrays and math on whole arrays.
What NVIDIA says (2)
“What is NumPy NumPy is a powerful, well-optimized, free open-source library for the Python programming language, adding support for large, multi-dimensional arrays (also called matrices or tensors).”
“As the core library for scientific computing, NumPy is the base for libraries such as Pandas, Scikit-learn , and SciPy .”
- NVIDIA AI Workbench
An NVIDIA tool that manages data science projects in containers with Git-managed code and environments.
What NVIDIA says (2)
“AI Workbench provides reproducibility by managing software, containers and Git repositories.”
“Versioned configuration files let you define and adapt environments to machines and users.”
- nvidia-smi
A command-line tool that lists NVIDIA GPUs and reports their status and use.
What NVIDIA says (1)
“nvidia-smi (also NVSMI) provides monitoring and management capabilities for each of NVIDIA's Tesla, Quadro, GRID and GeForce devices”
O
- One-hot encoding
Turning a category column into one 0/1 column per category.
What NVIDIA says (2)
“cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”
“However, dropping one category breaks the symmetry of the original representation and can therefore induce a bias in downstream models, for instance for penalized linear classification or regression models.”
- ONNX
A portable model file format that runs without the original training library.
What NVIDIA says (2)
“Use as_sklearn() to convert the cuML model to a scikit-learn estimator, then pass it to skl2onnx.convert_sklearn() .”
“The resulting .onnx file can be loaded with ONNX Runtime for inference on both CPU and GPU, with no cuML dependency at inference time.”
- Optuna
A framework that automates hyperparameter search and can run trials in parallel.
What NVIDIA says (2)
“Optuna is a lightweight framework for automatic hyperparameter optimization.”
“By simply wrapping the objective function with Optuna, we can perform a parallel-distributed HPO search over a search space as we’ll see in this notebook.”
- Overfitting
When a model learns noise in the training data and does poorly on new data.
What NVIDIA says (2)
“Overfitting means that the model may look very good on the training set but generalises poorly to new data that it has not seen before.”
“The presence of a large number of trees also reduces the problem of overfitting, which occurs when a model incorporates too much “noise” in the training data and makes poor decisions as a result.”
P
- p-value
The chance of seeing a result this extreme if there were really no effect; small values suggest a real effect.
What NVIDIA says (1)
“The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”
- PageRank
An importance score for each node, higher when important nodes link to it.
What NVIDIA says (1)
“Find the PageRank score for every vertex in a graph.”
- pandas
A Python library for working with tables of data, called DataFrames.
What NVIDIA says (2)
“A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable, and each row contains a set of values corresponding to those variables.”
“With its support for structured data formats like tables, matrices, and time series, the pandas Python API provides tools to process messy or raw datasets into clean, structured formats ready for analysis.”
- Parallel processing
Splitting a job into many pieces that run at the same time.
What NVIDIA says (2)
“By contrast, GPUs break complex problems into thousands or millions of separate tasks and work them out at once.”
“Dask is an open-source library designed to provide parallelism to the existing Python stack.”
- Parquet
A columnar file format that stores data by column, so tools can read only the columns they need.
What NVIDIA says (2)
“columns list, default None If not None, only these columns will be read.”
“If not None, specifies a filter predicate used to filter out row groups using statistics stored for each row group as Parquet metadata.”
- Partitioned dataset
A dataset split into folders by column values, such as one folder per year.
What NVIDIA says (1)
“partition_cols list, optional, default None Column names by which to partition the dataset Columns are partitioned in the order they are given”
- Personal identifiable information
Data that can identify a person, such as a name or phone number.
What NVIDIA says (1)
“Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”
- Pinned memory
Host memory the operating system cannot move, so the GPU can copy from it directly and faster.
What NVIDIA says (2)
“Higher bandwidth is possible between the host and the device when using page-locked (or “pinned”) memory.”
“The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”
- pip
Python's standard package installer; it cannot install system-level CUDA libraries.
What NVIDIA says (1)
“Since pip cannot install system-level CUDA libraries, CuPy expects these libraries to already be present in the system environment.”
- Plotly Dash
A Python framework for building interactive web dashboards.
What NVIDIA says (2)
“Plotly Dash enables data scientists to recast complex data and machine learning workflows as more accessible web applications.”
“The use of Plotly’s Dash, RAPIDS, and Data shader allows users to build viz dashboards that both render datasets of 300 million+ rows and remain highly interactive without the need for precomputed aggregations.”
- Polars GPU engine
A cuDF-powered engine that runs Polars lazy queries on the GPU and falls back to the CPU when needed.
What NVIDIA says (2)
“cuDF provides GPU-accelerated execution engines for Python users of the Polars Lazy API.”
“If it is not, the execution transparently falls back to the standard Polars engine and runs on the CPU.”
- Precision
Of the cases the model flagged as positive, the share that really are.
What NVIDIA says (1)
“The precision is the ratio tp / (tp + fp) where tp is the number of true positives and fp the number of false positives.”
- Prediction drift
A change in the distribution of a model's outputs, watched when true labels are not available.
What NVIDIA says (2)
“Prediction drift: When it is not possible to acquire ground truth labels, predictions must be monitored.”
“If there is a drastic change in the distribution of predictions, something has potentially gone wrong.”
- Prefect
A workflow orchestrator that chains pipeline stages, tracks runs, schedules them and handles failures.
What NVIDIA says (2)
“Prefect orchestrates the pipeline stages, tracks runs, and enables scheduling”
“Orchestration: Prefect flows chain the stages together, track runs, and handle failures”
- Principal component analysis
A method that combines features into a few linear components that keep the most variance.
What NVIDIA says (1)
“PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”
R
- R-squared
The share of the variation in the target that a regression model explains.
What NVIDIA says (2)
“R-squared (R²), also known as the coefficient of determination, represents the proportion of variance explained by a model.”
“Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”
- Random forest
Many decision trees trained on random samples whose votes are combined.
What NVIDIA says (2)
“Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”
“Whereas decision trees are based upon a fixed set of features, and often overfit, randomness is critical to the success of the forest.”
- Random search
Trying randomly sampled hyperparameter combinations instead of all of them.
What NVIDIA says (1)
“Random Search replaces the exhaustive nature of the search from before with a random selection of parameters over the specified space.”
- RAPIDS
NVIDIA's open-source GPU libraries for data science, covering loading, ETL, training and analysis.
What NVIDIA says (2)
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
“RAPIDS aims to accelerate the entire data science pipeline including data loading, ETL, model training, and inference.”
- RAPIDS Accelerator for Apache Spark
A plugin that runs Apache Spark jobs on GPUs without code changes.
What NVIDIA says (2)
“Run your existing Apache Spark applications with no code change.”
“The optimization came from using the RAPIDS accelerator for Apache Spark, an open-source library that enables GPU-accelerated ETL and feature engineering.”
- Recall
Of the real positive cases, the share the model found.
What NVIDIA says (2)
“The recall is the ratio tp / (tp + fn) where tp is the number of true positives and fn the number of false negatives.”
“The recall is intuitively the ability of the classifier to find all the positive samples.”
- Regression
Predicting a number, such as a price.
What NVIDIA says (1)
“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”
- Regularization
A penalty on model complexity that helps prevent overfitting.
What NVIDIA says (2)
“Ridge extends LinearRegression by providing L2 regularization on the coefficients when predicting response y with a linear combination of the predictors in X.”
“Without these regularisation terms, gradient boosted models can quickly become large and overfit to noise present in the training data.”
- Repository
A project folder plus its full Git history.
What NVIDIA says (1)
“Key Concepts # Project A Git repository under management by AI Workbench.”
- Resampling
Converting a time series to a new regular frequency, such as one value per minute.
What NVIDIA says (2)
“Parameters : rule: str The offset string representing the frequency to use.”
“First, we create a time series with 1 minute intervals”
- Rolling mean
The average over a sliding window of recent rows, used to smooth noise.
What NVIDIA says (2)
“Parameters : window int, offset or a BaseIndexer subclass Size of the window, i.e., the number of observations used to calculate the statistic.”
“Rolling window statistics are statistics (e.g. mean, standard deviation) over a time duration in the past.”
S
- Sampling
Taking a random subset of rows, for example for quick exploration.
What NVIDIA says (2)
“This function will always produce the same sample given an identical random_state .”
“frac float, optional Fraction of axis items to return.”
- scikit-learn
A popular CPU machine learning library for Python with a standard fit and predict API.
What NVIDIA says (1)
“cuML estimators look and feel just like scikit-learn estimators .”
- SHAP
A method that explains a single prediction by how much each feature pushed it.
What NVIDIA says (1)
“cuML’s SHAP based explainers accelerate the algorithmic part of SHAP.”
- Standard scaling
Rescaling a numeric feature to mean 0 and standard deviation 1.
What NVIDIA says (2)
“Standardize features by removing the mean and scaling to unit variance”
“Mean and standard deviation are then stored to be used on later data using transform() .”
- Stratified split
A train-test split that keeps the same class ratio in both parts.
What NVIDIA says (1)
“If not None, data is split in a stratified fashion, using this as the class labels.”
- Supervised learning
Learning from examples that come with the correct answer (labels).
What NVIDIA says (2)
“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”
“Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”
- Synthetic data
Generated data that mimics real data, used to fill gaps or protect privacy.
What NVIDIA says (2)
“Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”
“Synthetically generated data can augment existing data for larger, more representative datasets.”
T
- Target encoding
Replacing each category with the average target value for that category.
What NVIDIA says (2)
“The input data is grouped by the columns Xs and the aggregated mean value of Y of each group is calculated to replace each value of Xs .”
“Several optimizations are applied to prevent label leakage and parallelize the execution.”
- Test set
Data held back from training and used once to estimate performance on new data.
What NVIDIA says (1)
“test_size float or int, default=None If float, should be between 0.0 and 1.0 and represent the proportion of the dataset to include”
- Time series
Data points recorded in time order.
What NVIDIA says (2)
“Great care must be taken when defining cross-validation folds for time-series data.”
“First, we create a time series with 1 minute intervals”
U
- UMAP
A non-linear method that maps high-dimensional data to two or three dimensions for plotting.
What NVIDIA says (2)
“UMAP is a popular dimension reduction algorithm used in fields like bioinformatics, NLP topic modeling, and ML preprocessing.”
“It works by creating a k-nearest neighbors (k-NN) graph, which is known in literature as an all-neighbors graph, to build a fuzzy topological representation of the data, which is used to embed high-dimensional data into lower dimensions.”
- Underfitting
When a model is too simple to capture the pattern, even on training data.
What NVIDIA says (2)
“Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”
“Random forest bagging minimizes the variance and overfitting, while GBDT boosting reduces the bias and underfitting.”
- Unsupervised learning
Finding patterns in data that has no labels.
What NVIDIA says (2)
“Unsupervised learning algorithms attempt to ‘learn’ patterns in unlabeled data sets, discovering similarities, or regularities.”
“Common unsupervised tasks include clustering and association.”
W
- Weights & Biases
A platform for tracking, debugging and reproducing ML experiments.
What NVIDIA says (1)
“Weights & Biases : W&B’s tools help many ML users build better models faster, debugging and reproducing their models with just a few lines of code”
X
- XGBoost
A scalable library for gradient-boosted decision trees that can train on GPUs.
What NVIDIA says (1)
“XGBoost , which stands for Extreme Gradient Boosting, is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”