NCA-AIIO glossary

The official terms you will meet on the exam and in the field. Each has a one-sentence plain definition and the NVIDIA quote it is based on.

A

Accelerated computing

Using specialized hardware such as GPUs to speed up demanding work through parallel processing.

Objectives: 1.8, 1.4

What NVIDIA says (2)

“A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”

— CUDA Programming Guide: Introduction

“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”

— What Is Accelerated Computing?

Aisle containment

Enclosing the hot or cold aisle so hot exhaust air cannot mix back into server inlets.

Objectives: 2.3, 2.6

What NVIDIA says (1)

“In either case, the main benefit of aisle containment is the prevention of air recirculation from the hot aisle to the cold aisle”

— DGX SuperPOD Data Center Design (H100): Cooling

Artificial intelligence (AI)

Software that performs tasks that normally need human intelligence; machine learning and deep learning are subsets of it.

Objectives: 1.3

What NVIDIA says (1)

“As a subset of AI, machine learning in its most elemental form uses algorithms to parse data, learn from it, and then make predictions or determinations about something in the real world.”

— What is Machine Learning and Why Does It Matter?

B

Bandwidth

How much data a link or memory can move per second, for example in GB/s.

Objectives: 2.5

What NVIDIA says (1)

“640 GB of aggregated HBM3 memory with 24 TB/s of aggregate memory bandwidth”

— DGX SuperPOD H100 Reference Architecture: Components

Batch (batch size)

The group of samples a model processes in one step; batch size is how many samples are in that group.

Objectives: 1.2

What NVIDIA says (1)

“Mixed precision models use less memory than FP32, so it is possible to increase the batch size when running with AMP.”

— NVIDIA Mixed Precision Training Guide

C

CapEx and OpEx (capital and operational expenditure)

CapEx is money spent up front on assets; OpEx is ongoing spending such as a cloud bill.

Objectives: 2.4

What NVIDIA says (1)

“Cloud-based solutions offer a cost-effective way to start AI initiatives by reducing acquisition costs and shifting capital expenditures (CapEx) to operational expenditures (OpEx).”

— What Is AI Infrastructure? (NVIDIA Glossary)

Central processing unit (CPU)

The general-purpose processor in a server: a few fast cores with large caches, built to run a few threads quickly.

Objectives: 1.4

What NVIDIA says (3)

“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”

— CUDA Programming Guide: Introduction

“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”

— CUDA Programming Guide: Introduction

“Architecturally, the CPU is composed of just a few cores with lots of cache memory that can handle a few software threads at a time.”

— What's the Difference Between a CPU and a GPU?

Checkpoint

A saved copy of training state so a job can resume after a failure.

Objectives: 2.5

What NVIDIA says (1)

“As DL models grow in size and time-to-train, writing checkpoints is necessary for fault tolerance.”

— DGX SuperPOD H100 Reference Architecture: Storage Architecture

Container

A package of an application together with its software dependencies (libraries, configuration, tools) that runs the same on any host.

Objectives: 1.1, 3.2

What NVIDIA says (1)

“All of them also include the needed GPU libraries, configuration files, and tools to rebuild the container.”

— NVIDIA Deep Learning Frameworks User Guide (NGC containers)

CUDA

NVIDIA's parallel computing platform; CUDA-X libraries built on it accelerate AI and HPC.

Objectives: 1.1

What NVIDIA says (3)

“to enable any computational workload to use the throughput capability of GPUs independent of graphics APIs.”

— CUDA Programming Guide: Introduction

“Libraries like cuBLAS, cuFFT, cuDNN, and CUTLASS are just a few examples of libraries that help developers avoid reimplementing well-established algorithms.”

— CUDA Programming Guide: Introduction

“NVIDIA CUDA-X™, built on CUDA, is a collection of libraries that deliver dramatically higher performance across application domains, including AI and HPC.”

— CUDA Platform for Accelerated Computing

cuDNN (CUDA Deep Neural Network library)

A GPU library of tuned building blocks, such as convolutions, for deep neural networks.

Objectives: 1.1

What NVIDIA says (1)

“The NVIDIA CUDA Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives for deep neural networks.”

— NVIDIA cuDNN Documentation

D

Data processing unit (DPU)

A system on a chip with CPU cores, a fast network interface and acceleration engines that moves and processes data center traffic.

Objectives: 2.10

What NVIDIA says (2)

“BlueField-3 DPUs offload, accelerate, and isolate software-defined networking, storage, security, and management functions, significantly enhancing data center performance, efficiency, and security.”

— NVIDIA BlueField-3 Networking Platform User Guide: Introduction

“Specialists in moving data in data centers, DPUs, or data processing units, are a new class of programmable processor and will join CPUs and GPUs as one of the three pillars of computing.”

— What Is a DPU? (NVIDIA Blog)

DCGM (Data Center GPU Manager)

NVIDIA tooling for GPU health monitoring, diagnostics, alerts and policies across a cluster.

Objectives: 3.3, 3.1

What NVIDIA says (3)

“Active Health Checks (GPU subsystems)”

— DCGM Feature Overview

“GPU Diagnostics (Diagnostic Levels - 1, 2, 3)”

— DCGM Feature Overview

“It includes active health monitoring, comprehensive diagnostics, system alerts, and governance policies including power and clock management.”

— NVIDIA DCGM (developer page)

DCGM Exporter

A service that publishes DCGM GPU metrics on an HTTP /metrics endpoint for Prometheus.

Objectives: 3.3

What NVIDIA says (1)

“DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”

— NVIDIA DCGM Exporter

Deep learning (DL)

A subset of machine learning that uses neural networks with many layers to learn patterns directly from data such as images or text.

Objectives: 1.3, 1.4

What NVIDIA says (1)

“Deep learning is a subset of machine learning, with the difference that DL algorithms can automatically learn representations from data such as images, video, or text, without introducing human domain knowledge.”

— What Is Deep Learning and Why Does It Matter?

DGX SuperPOD

NVIDIA's reference design for an AI cluster made of DGX systems, InfiniBand and Ethernet networking, management nodes and storage.

Objectives: 2.2, 2.5

What NVIDIA says (1)

“The DGX SuperPOD architecture is a combination of DGX systems, InfiniBand and Ethernet networking, management nodes, and storage.”

— DGX SuperPOD H100 Reference Architecture: Architecture

DGX system (DGX)

NVIDIA's integrated GPU server; a DGX H100 system has eight GPUs.

Objectives: 2.3

What NVIDIA says (1)

“The DGX H100 system, which is the fourth-generation NVIDIA DGX system, delivers AI excellence in an eight GPU configuration.”

— DGX SuperPOD H100 Reference Architecture: Components

DOCA

The software framework (SDK and runtime) for building services on BlueField DPUs and SuperNICs.

Objectives: 2.10

What NVIDIA says (1)

“DOCA contains a runtime and development environment, including libraries and drivers for device management and programmability”

— NVIDIA DOCA Overview

E

ECC (error-correcting code)

Memory protection that fixes single-bit errors and detects double-bit errors.

Objectives: 3.1, 3.3

What NVIDIA says (1)

“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”

— DCGM Field IDs

F

Fabric (network fabric)

The network that connects the servers: the switches, cables and adapters together.

Objectives: 2.8

What NVIDIA says (1)

“An InfiniBand fabric is composed of switches and channel adapter (HCA/TCA) devices.”

— MLNX_OFED: InfiniBand Fabric Utilities

G

Generative AI

AI models that learn patterns in existing data and use them to create new content such as text, images or code.

Objectives: 1.3, 1.5

What NVIDIA says (1)

“Generative AI models use neural networks to identify the patterns and structures within existing data to generate new and original content.”

— What is Generative AI and How Does it Work?

GPU Operator

A Kubernetes operator that installs and manages the NVIDIA driver, device plugin, container toolkit and monitoring.

Objectives: 3.2

What NVIDIA says (1)

“The NVIDIA GPU Operator uses the operator framework within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU.”

— About the NVIDIA GPU Operator

GPUDirect RDMA

A direct path between GPU memory and a peer device such as a network adapter over PCI Express.

Objectives: 2.8

What NVIDIA says (1)

“GPUDirect RDMA is a technology introduced in Kepler-class GPUs and CUDA 5.0 that enables a direct path for data exchange between the GPU and a third-party peer device using standard features of PCI Express.”

— GPUDirect RDMA

GPUDirect Storage (GDS)

A direct path that moves data between storage and GPU memory without a copy through CPU memory.

Objectives: 2.5

What NVIDIA says (1)

“GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which avoids a bounce buffer through the CPU.”

— GPUDirect Storage Overview

Gradient

The correction signal computed in training that says how much, and in which direction, to adjust each weight.

Objectives: 1.2, 2.6

What NVIDIA says (1)

“Distributed Data Parallelism (DDP) keeps the model copies consistent by synchronizing parameter gradients across data-parallel GPUs before each parameter update.”

— NVIDIA Megatron Bridge: Parallelisms Guide

Graphics processing unit (GPU)

A processor with hundreds or thousands of cores that runs many threads in parallel for high total throughput.

Objectives: 1.8

What NVIDIA says (2)

“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”

— CUDA Programming Guide: Introduction

“In contrast, a GPU is composed of hundreds of cores that can handle thousands of threads simultaneously.”

— What's the Difference Between a CPU and a GPU?

I

Inference

Using a trained model to make predictions or generate outputs on new data.

Objectives: 1.2

What NVIDIA says (1)

“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”

— What's the Difference Between Deep Learning Training and Inference?

InfiniBand (IB)

A high-performance, low-latency, RDMA-capable network used to connect GPU servers and storage.

Objectives: 2.8, 2.9

What NVIDIA says (1)

“InfiniBand is a high-performance, low latency, RDMA capable networking technology”

— DGX SuperPOD H100 Reference Architecture: Components

K

Kubernetes (K8s)

An open-source system that orchestrates containers: it deploys, scales and manages them across a cluster.

Objectives: 3.2

What NVIDIA says (1)

“Kubernetes (K8s) is an open-source system for automating deployment, scaling, and management of containerized applications.”

— NVIDIA DeepOps: GPU cluster deployment

L

Large language model (LLM)

A very large neural network trained on huge text datasets that can generate human-like text.

Objectives: 1.3, 1.5

What NVIDIA says (1)

“However, large language models, which are trained on internet-scale datasets with hundreds of billions of parameters, have now unlocked an AI model’s ability to generate human-like content.”

— What are Large Language Models?

Latency

How long one request or one transfer takes from start to finish.

Objectives: 1.2

What NVIDIA says (1)

“TensorRT includes inference compilers, runtimes, and model optimizations that deliver low latency and high throughput for production applications.”

— NVIDIA TensorRT (developer page)

Lossless Ethernet (lossless)

Ethernet configured to pause traffic instead of dropping packets when buffers fill, which RoCE needs.

Objectives: 2.8

What NVIDIA says (2)

“RDMA over Converged Ethernet (RoCE) is a mechanism to provide this efficient data transfer with very low latencies on lossless Ethernet networks.”

— MLNX_OFED: RDMA over Converged Ethernet (RoCE)

“For example, PFC can provide lossless service for the RoCE traffic and best-effort service for the standard Ethernet traffic.”

— MLNX_OFED: Flow Control (PFC)

M

Machine learning (ML)

A subset of AI where algorithms learn from data to make predictions instead of following hand-written rules.

Objectives: 1.3

What NVIDIA says (1)

“As a subset of AI, machine learning in its most elemental form uses algorithms to parse data, learn from it, and then make predictions or determinations about something in the real world.”

— What is Machine Learning and Why Does It Matter?

Mixed precision

Training with lower-precision number formats where safe, which saves memory and runs math faster on Tensor Cores.

Objectives: 1.2, 1.8

What NVIDIA says (1)

“Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”

— NVIDIA Mixed Precision Training Guide

MLOps (machine learning operations)

Practices for building, deploying and maintaining ML models in production, extending DevOps.

Objectives: 1.7

What NVIDIA says (1)

“short for machine learning operations, is a set of practices and principles that aims to streamline the development, deployment, and maintenance of machine learning (ML) models in production environments.”

— What Is MLOps? (NVIDIA Glossary)

Multi-Instance GPU (MIG)

Splitting one GPU into up to seven isolated instances, each with its own compute and memory.

Objectives: 3.4

What NVIDIA says (1)

“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”

— MIG User Guide: Introduction

N

N+1 redundancy

Providing the N power sources you need plus one spare, so one can fail without an outage.

Objectives: 2.6

What NVIDIA says (1)

“Due to this requirement, the data center must minimally provide N+1 power, where N equals two power sources.”

— DGX SuperPOD Data Center Design (H100): Electrical

NCCL (NVIDIA Collective Communications Library)

The library that moves data between GPUs inside and across nodes, with operations such as all-reduce.

Objectives: 1.1, 2.7

What NVIDIA says (2)

“is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”

— Overview of NCCL — NCCL 2.32.3 documentation

“NCCL implements both collective communication and point-to-point send/receive primitives.”

— Overview of NCCL — NCCL 2.32.3 documentation

NGC catalog

NVIDIA's catalog of GPU-optimized containers, pretrained models and SDKs.

Objectives: 1.1

What NVIDIA says (1)

“The NGC Catalog consists of containers, pretrained models, Helm charts for Kubernetes deployments, and industry-specific AI toolkits with software development kits (SDKs).”

— NGC Catalog User Guide

NVIDIA NIM

Prebuilt, optimized inference microservices for deploying AI models on NVIDIA GPUs anywhere.

Objectives: 1.6

What NVIDIA says (2)

“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”

— NVIDIA NIM Documentation

“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”

— NVIDIA NIM Microservices

NVIDIA's direct GPU-to-GPU interconnect inside a server.

Objectives: 2.9, 2.5

What NVIDIA says (1)

“is a direct GPU-to-GPU interconnect that scales multi-GPU input/output (IO) in the server.”

— NVIDIA Fabric Manager User Guide

A switch chip that connects many NVLinks so all GPUs can talk to each other at full NVLink speed.

Objectives: 2.9

What NVIDIA says (2)

“which connects multiple NVLinks to provide all-to-all GPU communication at the total NVLink speed.”

— NVIDIA Fabric Manager User Guide

“The NVIDIA NVLink Switch chips connect multiple NVLinks to provide all-to-all GPU communication at full NVLink speed across the entire rack.”

— NVIDIA NVLink and NVLink Switch

O

Oversubscription

When more bandwidth (or demand) can enter a layer than its capacity can serve.

Objectives: 2.8, 2.3

What NVIDIA says (1)

“When there is a delta between the capacity of a resource in the data center and the demand for that resource, the resource is said to be oversubscribed.”

— DGX SuperPOD Data Center Design (H100): Cooling

P

PCI Express (PCIe)

The standard interconnect inside a server that links the CPU to GPUs, network adapters and storage.

Objectives: 2.5

What NVIDIA says (2)

“It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”

— Overview of NCCL — NCCL 2.32.3 documentation

“These routines are optimized to achieve high bandwidth and low latency over PCIe,NVIDIA NVLink™, and other high-speed interconnects within a node and over NVIDIA networking across nodes.”

— NVIDIA Collective Communications Library (developer page)

Pretrained model

A model already trained on a large dataset that you can use as is or fine-tune for your task.

Objectives: 1.4, 1.7

What NVIDIA says (1)

“A pretrained AI model is a deep learning model that’s trained on large datasets to accomplish a specific task, and it can be used as is or customized to suit application requirements across multiple industries.”

— What Is a Pretrained AI Model?

Prometheus

An open-source monitoring system that collects metrics, such as DCGM Exporter GPU metrics, from /metrics endpoints.

Objectives: 3.3

What NVIDIA says (1)

“DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”

— NVIDIA DCGM Exporter

R

Rack density

How many systems are placed in one rack; it is limited by the power and cooling available per rack.

Objectives: 2.2, 2.3

What NVIDIA says (1)

“However, rack densities can be customized to fit within the available power and cooling capacities at the data center.”

— DGX SuperPOD Data Center Design (H100): Planning

Rail-optimized network

A compute fabric where each same-numbered port on every node connects to its own set of leaf switches, keeping that traffic one hop apart.

Objectives: 2.7

What NVIDIA says (1)

“Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”

— DGX SuperPOD H100 Reference Architecture: Network Fabrics

RDMA over Converged Ethernet (RoCE)

RDMA carried over Ethernet networks instead of InfiniBand.

Objectives: 2.8, 2.9

What NVIDIA says (1)

“RDMA over Converged Ethernet (RoCE) is a mechanism to provide this efficient data transfer with very low latencies on lossless Ethernet networks.”

— MLNX_OFED: RDMA over Converged Ethernet (RoCE)

Remote direct memory access (RDMA)

Moving data straight into another machine's memory without the CPU copying it, which cuts latency and CPU load.

Objectives: 2.8

What NVIDIA says (1)

“Storage is provided over InfiniBand and leverages RDMA to provide maximum performance and minimize CPU overhead.”

— DGX SuperPOD H100 Reference Architecture: Components

S

Scalable unit (SU)

A repeatable building block of 32 DGX H100 systems used to grow a DGX SuperPOD.

Objectives: 2.2

What NVIDIA says (1)

“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”

— DGX SuperPOD H100 Reference Architecture: Abstract

SHARP (Scalable Hierarchical Aggregation and Reduction Protocol)

Switch-based in-network computing that performs collective operations so less data crosses the network.

Objectives: 2.9

What NVIDIA says (1)

“technology improves the performance of MPI and Machine Learning collective operation, by offloading collective operations from CPUs and GPUs to the network”

— NVIDIA SHARP Documentation: Introduction

Slurm

An open-source cluster scheduler that queues jobs and assigns them nodes and GPUs.

Objectives: 3.2

What NVIDIA says (1)

“Slurm is an open-source cluster resource management and job scheduling system”

— NVIDIA DeepOps: GPU cluster deployment

Spectrum-X Ethernet

NVIDIA's AI-optimized Ethernet platform of switches and SuperNICs.

Objectives: 2.9

What NVIDIA says (3)

“NVIDIA Spectrum-X is an AI-optimized Ethernet networking platform that combines NVIDIA Spectrum switches with the BlueField-3 SuperNIC”

— NVIDIA Network Operator: Spectrum-X Ethernet Networking Platform

“to deliver high-bandwidth, lossless RoCE for the GPU-to-GPU compute (east-west) network.”

— NVIDIA Network Operator: Spectrum-X Ethernet Networking Platform

“It combines purpose-built Ethernet switches and SuperNICs to deliver high effective bandwidth, low latency, and strong performance isolation for hyperscale AI infrastructure”

— NVIDIA Spectrum-X Ethernet

Streaming multiprocessor (SM)

The GPU compute unit that runs threads; a GPU has many SMs, and MIG divides them between instances.

Objectives: 3.4

What NVIDIA says (1)

“available GPU compute resources (including streaming multiprocessors or SMs, and GPU engines such as copy engines or decoders)”

— MIG User Guide: Introduction

Subnet Manager (SM)

The service that configures an InfiniBand fabric; one must be running at all times.

Objectives: 2.8

What NVIDIA says (1)

“All InfiniBand-compliant ULPs require a proper operation of a Subnet Manager (SM) running on the InfiniBand fabric, at all times.”

— MLNX_OFED Documentation: Introduction

T

Tensor Core

A specialized unit inside NVIDIA GPUs that speeds up the matrix math used in deep learning.

Objectives: 1.8

What NVIDIA says (1)

“Tensor Cores were introduced in the NVIDIA Volta™ GPU architecture to accelerate matrix multiply and accumulate operations for machine learning and scientific applications.”

— GPU Performance Background User's Guide

Throughput

How much total work finishes per second, for example requests or samples per second.

Objectives: 1.2

What NVIDIA says (1)

“they allow enterprises to optimize throughput and latency so they can meet their service level agreements across a variety of use cases.”

— What's the Difference Between Deep Learning Training and Inference?

Time-slicing

Sharing a GPU by letting workloads take turns, without the memory and fault isolation of MIG.

Objectives: 3.4

What NVIDIA says (1)

“Time-slicing trades the memory and fault-isolation that is provided by MIG for the ability to share a GPU by a larger number of users.”

— Time-Slicing GPUs in Kubernetes

Total cost of ownership (TCO)

The full cost of a system over its life, including compute, storage and maintenance.

Objectives: 2.4

What NVIDIA says (1)

“IT leaders should evaluate the total cost of ownership (TCO) over time and consider factors such as data storage, compute resources, and ongoing maintenance.”

— What Is AI Infrastructure? (NVIDIA Glossary)

Training

The process of feeding data through a model and adjusting its weights until its outputs are accurate.

Objectives: 1.2

What NVIDIA says (1)

“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”

— What's the Difference Between Deep Learning Training and Inference?

Transformer

A neural network design that uses attention to relate every part of an input to every other part, and runs well in parallel.

Objectives: 1.3, 1.4

What NVIDIA says (1)

“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.”

— What Is a Transformer Model?

Triton Inference Server

Open-source software that serves models from many frameworks for inference.

Objectives: 1.6

What NVIDIA says (1)

“Triton Inference Server is an open source inference serving software that streamlines AI inferencing.”

— NVIDIA Triton Inference Server

V

Virtual GPU (vGPU)

NVIDIA software that creates virtual GPUs shared across virtual machines.

Objectives: 3.4

What NVIDIA says (1)

“NVIDIA virtual GPU software enables multiple virtual machines (VMs) to have simultaneous, direct access to a single physical GPU”

— NVIDIA Virtual GPU Software Documentation

X

Xid error

An error report the NVIDIA driver writes to the kernel log, identified by a number such as 48 or 79.

Objectives: 3.1

What NVIDIA says (1)

“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

— NVIDIA Xid Catalog