Essential AI Knowledge
38% of the NCA-AIIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Essential AI Knowledge · AI Infrastructure · AI Operations
1.1 The NVIDIA software stack
The layers between the GPU and your AI application: driver, CUDA, libraries, containers and serving.
Key points
The CUDA toolkit is what developers build with. The driver is what runs on the server. They are versioned separately. NVIDIA's CUDA Compatibility guide defines supported ways to run newer CUDA software on existing driver installations, within documented limits. Forward compatibility uses the cuda-compat package so applications built with a newer toolkit can run on older base drivers, subject to platform and GPU support.
What NVIDIA says (2)
“CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”
“It also includes Forward Compatibility , which uses the cuda-compat-<major>-<minor> package to allow applications built with a newer toolkit to run on older base drivers across major release families, subject to platform and GPU support.”
A collective is a communication step that involves every GPU in a job. All-reduce is a collective that sums values, such as gradients, across all GPUs and gives every GPU the result. NCCL (NVIDIA Collective Communications Library) implements multi-GPU and multi-node communication primitives. It provides all-gather, all-reduce, broadcast, reduce, reduce-scatter, and point-to-point send and receive.
What NVIDIA says (3)
“is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”
“NCCL implements both collective communication and point-to-point send/receive primitives.”
“NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”
CUDA is NVIDIA's platform for GPU computing. CUDA-X is the collection of libraries built on it. cuDNN (CUDA Deep Neural Network library) is a GPU-accelerated library of primitives for deep neural networks. A primitive is a basic building block that frameworks such as PyTorch call under the hood. NVIDIA lists attention, convolution, matrix multiplication, normalization and pooling among the operations it tunes.
What NVIDIA says (5)
“The NVIDIA CUDA Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives for deep neural networks.”
“It provides highly tuned implementations of operations arising frequently in deep neural network (DNN) applications: Scaled dot-product attention Convolution, including cross-correlation Matrix multiplication Normalizations, softmax, and pooling”
“to enable any computational workload to use the throughput capability of GPUs independent of graphics APIs.”
“Libraries like cuBLAS, cuFFT, cuDNN, and CUTLASS are just a few examples of libraries that help developers avoid reimplementing well-established algorithms.”
“NVIDIA CUDA-X™, built on CUDA, is a collection of libraries that deliver dramatically higher performance across application domains, including AI and HPC.”
RAPIDS is NVIDIA's set of CUDA-X data science libraries. cuDF accelerates fundamental DataFrame operations, the table operations pandas users run. cuML is a GPU-accelerated machine learning library. NVIDIA says zero-code-change APIs accelerate popular tools like pandas and scikit-learn.
What NVIDIA says (4)
“is a GPU-accelerated library for tabular data processing.”
“NVIDIA cuML is a suite of fast, GPU-accelerated machine learning algorithms designed for data science and analytical tasks.”
“a zero-code change accelerator, cudf.pandas , for existing pandas code.”
“automatically accelerate existing code with zero code changes.”
A container packages an application with everything it needs to run, so it behaves the same everywhere. The NGC catalog gives access to GPU-accelerated software: performance-optimized containers, pretrained AI models and industry-specific SDKs (software development kits). NVIDIA says these can be deployed on premises, in the cloud or at the edge.
What NVIDIA says (2)
“The NGC Catalog consists of containers, pretrained models, Helm charts for Kubernetes deployments, and industry-specific AI toolkits with software development kits (SDKs).”
“The NGC catalog provides access to GPU-accelerated software that speeds up end-to-end workflows with performance-optimized containers, pretrained AI models, and industry-specific SDKs that can be deployed on premises, in the cloud, or at the edge.”
Containers normally cannot see a server's GPUs. The NVIDIA Container Toolkit provides the runtime pieces that expose GPUs to containers. NVIDIA lists components such as the NVIDIA Container Runtime and the nvidia-ctk command-line tool. The GPU Operator installs it, together with the driver, on Kubernetes nodes.
What NVIDIA says (2)
“It currently includes: The NVIDIA Container Runtime ( nvidia-container-runtime ) The NVIDIA Container Toolkit CLI ( nvidia-ctk )”
“These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the NVIDIA Container Toolkit”
Key terms: CUDA cuDNN NCCL NGC catalog Container
Try it: CUDA Stack Verification NGC Container Flow NVIDIA AI Software Stack
1.2 Training vs inference
What training and inference each do, and why they need different resources.
Key points
Training is how a model learns. Data passes through the network, the error is propagated back through its layers, and the model adjusts its weights and tries again. A weight is one of the numbers inside the model that training tunes. Inference is when a trained model makes predictions on new data. NVIDIA notes inference cannot happen without training.
What NVIDIA says (3)
“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”
“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”
“Inference can’t happen without training.”
Precision is how many bits a number uses. FP16 (half precision) uses 16 bits; FP32 (single precision) uses 32. Mixed precision does most operations in half precision while keeping critical information in single precision. NVIDIA lists three benefits: less memory, less memory bandwidth, and much faster math, especially on GPUs with Tensor Cores.
What NVIDIA says (3)
“First, they require less memory, enabling the training and deployment of larger neural networks.”
“Second, they require less memory bandwidth which speeds up data transfer operations.”
“Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”
NVIDIA explains that training repeats a loop: the error is propagated back through the layers, the model adjusts its weights and tries again. Over many iterations it converges on the correct weights. NVIDIA calls training 'a monster when it comes to consuming compute.'
What NVIDIA says (2)
“Over many iterations, it converges on the correct weights, enabling it to reliably suggest accurate, useful code completions or translations.”
“The challenge is, it’s also a monster when it comes to consuming compute.”
Latency is the time from a request to its response. Throughput is how much work is completed per unit of time. NVIDIA says inference applies learned knowledge in real time. Low-latency, high-performance inference makes real-time interaction possible. NVIDIA adds that inference optimization lets enterprises tune throughput and latency to meet their service-level agreements (SLAs).
What NVIDIA says (2)
“This is only possible with low-latency, high-performance inference.”
“That’s why inference optimization techniques are so critical: they allow enterprises to optimize throughput and latency so they can meet their service level agreements across a variety of use cases.”
A large language model (LLM) generates text one token at a time. To avoid recomputing earlier tokens, it stores their attention keys and values in a KV (key-value) cache. NVIDIA says the two main contributors to GPU memory for LLM inference are the model weights and the KV cache. With batching, each request's KV cache is allocated separately, which can use a lot of memory.
What NVIDIA says (2)
“In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”
“With batching, the KV cache of each of the requests in the batch must still be allocated separately, and can have a large memory footprint.”
Key terms: Training Inference Mixed precision Latency Throughput Batch Gradient
Try it: Distributed Training (DDP) Training vs Inference Serving
1.3 AI, machine learning and deep learning
How the three terms nest inside each other, and where generative AI and LLMs fit.
Key points
A neural network is a model built from layers of simple connected units. NVIDIA says 'deep' represents the many layers of algorithms, or neural networks, used to recognize patterns in data.
What NVIDIA says (1)
“The word "deep" in deep learning represents the many layers of algorithms, or neural networks, that are used to recognize patterns in data.”
Artificial intelligence (AI) is the broad field. NVIDIA describes machine learning as a subset of AI that uses algorithms to parse data, learn from it, and make predictions. Deep learning is a subset of machine learning that can learn representations from data such as images, video or text without human domain knowledge.
What NVIDIA says (2)
“As a subset of AI, machine learning in its most elemental form uses algorithms to parse data, learn from it, and then make predictions or determinations about something in the real world.”
“Deep learning is a subset of machine learning, with the difference that DL algorithms can automatically learn representations from data such as images, video, or text, without introducing human domain knowledge.”
Generative AI models use neural networks to identify patterns and structures in existing data. They then generate new and original content. NVIDIA says inputs and outputs can include text, image, audio, video and code.
What NVIDIA says (2)
“Generative AI models use neural networks to identify the patterns and structures within existing data to generate new and original content.”
“Generative AI models can take inputs such as text, image, audio, video, and code and generate new content into any of the modalities mentioned.”
A parameter is one of the learned numbers inside a model. NVIDIA says LLMs are trained on internet-scale datasets with hundreds of billions of parameters. That scale has unlocked the ability to generate human-like content.
What NVIDIA says (1)
“However, large language models, which are trained on internet-scale datasets with hundreds of billions of parameters, have now unlocked an AI model’s ability to generate human-like content.”
NVIDIA describes training as passing data through layered connections where each neuron assigns weights to its input. The weights are adjusted based on feedback about whether the output was right or wrong. The trained model is, in effect, that balanced set of weights.
What NVIDIA says (2)
“Training a neural network, unlike human learning, involves passing data through layered connections where each neuron assigns weights and adjusts based only on “right” or “wrong” feedback.”
“Now you have a data structure and all the weights in there have been balanced based on what it has learned as you sent the training data through.”
A transformer is the neural network design behind most modern LLMs. NVIDIA says transformers use attention, or self-attention, to detect how even distant elements in a series influence each other. The math transformers use lends itself to parallel processing, so they run fast on parallel hardware like GPUs.
What NVIDIA says (2)
“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.”
“In addition, the math that transformers use lends itself to parallel processing, so these models can run fast.”
Key terms: Artificial intelligence Machine learning Deep learning Generative AI Large language model Transformer
Try it: AI, ML & DL Foundations
1.4 Why AI took off
The factors behind the recent jump in AI: GPUs, transformers and pretrained models.
Key points
Parallel computing means doing many calculations at the same time. NVIDIA says the parallelism of deep learning maps naturally to GPUs. That gives a significant speedup over CPU-only training and made GPUs the platform of choice for large neural networks. NVIDIA notes GPU-accelerated frameworks can cut training from days and weeks to hours and days.
What NVIDIA says (2)
“This parallelism maps naturally to GPUs , providing a significant computation speedup over CPU-only training and making them the platform of choice for training large, complex neural network-based systems.”
“researchers and data scientists can significantly speed up deep learning training that could otherwise take days and weeks to just hours and days.”
Labeled data has the right answer attached to each example, which takes people time and money to create. Self-supervised learning finds its training signal in the data itself. NVIDIA says that before transformers, users had to train with large, labeled datasets that were costly and time-consuming to produce. NVIDIA's CEO is quoted: transformers made self-supervised learning possible, and AI jumped to warp speed.
What NVIDIA says (2)
“Before transformers arrived, users had to train neural networks with large, labeled datasets that were costly and time-consuming to produce.”
“Transformers made self-supervised learning possible, and AI jumped to warp speed”
A pretrained model is a deep learning model trained on large datasets to do a specific task. Fine-tuning is extra training that adapts it to a narrower need. NVIDIA says a pretrained model can be used as is or customized to suit application requirements across multiple industries.
What NVIDIA says (3)
“A pretrained AI model is a deep learning model that’s trained on large datasets to accomplish a specific task, and it can be used as is or customized to suit application requirements across multiple industries.”
“You can use the pretrained models for inference or fine-tune them with transfer learning.”
“It can be used as is or further fine-tuned to fit an application’s specific needs.”
Key terms: Deep learning Transformer Pretrained model Accelerated computing Central processing unit
Try it: AI, ML & DL Foundations
1.5 AI use cases and industries
Where AI is used today and what kind of model each use case needs.
Key points
Computer vision is AI that interprets images and video. NVIDIA notes computer vision can help spot diseases quickly in medical imaging. Modern neural networks and GPU computing have led to its adoption across industries including healthcare, retail, manufacturing, transportation and financial services.
What NVIDIA says (2)
“They can also help spot diseases quickly in medical imaging and save lives .”
“have led to widespread adoption across industries like transportation , retail , manufacturing , healthcare , and financial services .”
NVIDIA defines a recommendation system as an AI algorithm, usually machine learning, that uses Big Data to suggest or recommend additional products to consumers. It learns preferences and past decisions from interaction data.
What NVIDIA says (2)
“A recommendation system is an artificial intelligence or AI algorithm, usually associated with machine learning , that uses Big Data to suggest or recommend additional products to consumers.”
“Recommender systems are trained to understand the preferences, previous decisions, and characteristics of people and products using data gathered about their interactions.”
NVIDIA groups LLM text use cases into generation (such as marketing content), summarization (such as meeting notes), translation (including text-to-code), classification (such as sentiment analysis) and chatbots (such as virtual assistants).
What NVIDIA says (1)
“Broadly, LLM use cases for text-based content can be divided up in the following manner: Generation (e.g., story writing, marketing content creation) Summarization (e.g., legal paraphrasing, meeting notes summarization) Translation (e.g., between languages, text-to-code) Classification (e.g., toxicity classification, sentiment analysis) Chatbot (e.g., open-domain Q+A, virtual assistants)”
NVIDIA says companies are increasingly data-driven. They use analytics and machine learning to recognize complex patterns, detect changes and make predictions that directly impact the bottom line.
What NVIDIA says (1)
“Companies are increasingly data-driven–sensing market and environment data, and using analytics and machine learning to recognize complex patterns, detect changes, and make predictions that directly impact the bottom line.”
Key terms: Generative AI Large language model
Try it: AI, ML & DL Foundations
1.6 NVIDIA solutions and what they are for
Which NVIDIA product solves which problem: NIM, Triton, NeMo, AI Enterprise, DGX and HGX.
Key points
A microservice is a small, self-contained service that other applications call over an API (application programming interface). NVIDIA NIM provides prebuilt, optimized inference microservices for deploying AI models on NVIDIA-accelerated infrastructure. NIM containers expose industry-standard APIs and are tuned for latency and throughput.
What NVIDIA says (3)
“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”
“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”
“expose industry-standard APIs for simple integration into AI applications, development frameworks, and workflows and optimize response latency and throughput for each combination of foundation model and GPU.”
NVIDIA describes the NeMo Framework as a scalable, cloud-native generative AI framework. It is built for researchers and developers working on LLMs, multimodal models and speech AI.
What NVIDIA says (1)
“NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI”
Inference serving means running trained models behind an endpoint that applications can call. Triton is open-source inference serving software. It deploys models from frameworks such as TensorRT, PyTorch and ONNX. Its features include concurrent model execution and dynamic batching, which groups requests to use the GPU better.
What NVIDIA says (3)
“Triton Inference Server is an open source inference serving software that streamlines AI inferencing.”
“Triton Inference Server enables teams to deploy any AI model from multiple deep learning and machine learning frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL, and more.”
“Concurrent model execution Dynamic batching”
NVIDIA AI Enterprise is NVIDIA's cloud-native software platform for production AI. NVIDIA says it increases reliability with extended-lifetime production branches and enterprise support. A production branch is a software release line kept stable and patched for a long time.
What NVIDIA says (1)
“Increase reliability with extended-lifetime production branches and enterprise support.”
cuDF is the RAPIDS library for DataFrame operations on the GPU. NVIDIA offers plug-in accelerators for DataFrame libraries like pandas with no code changes required.
What NVIDIA says (4)
“is a GPU-accelerated library for tabular data processing.”
“a zero-code change accelerator, cudf.pandas , for existing pandas code.”
“A Python library providing a GPU engine for Polars”
“accelerators for popular DataFrame libraries and SQL engines, like Polars, pandas, and Apache Spark with no code changes required.”
NVIDIA calls DGX a complete AI solution with full-stack, intelligent software. HGX brings together NVIDIA GPUs, CPUs, NVLink, networking and optimized software. NVIDIA says HGX is available as a single baseboard with eight GPUs, which server makers build systems around.
What NVIDIA says (3)
“Unlock productivity with the NVIDIA DGX platform, a complete AI solution with full-stack, intelligent software in NVIDIA Mission Control .”
“This server variant consists of one GPU baseboard with eight NVIDIA H100 GPUs and four NVSwitches.”
“NVIDIA HGX is available in a single baseboard with eight NVIDIA Rubin, NVIDIA Blackwell, or NVIDIA Blackwell Ultra SXMs.”
Key terms: NVIDIA NIM Triton Inference Server
Try it: NVIDIA AI Software Stack
1.7 The AI development and deployment life cycle
The software that supports each step from data to a model in production.
Key points
The AI life cycle runs from data preparation to model development, deployment and monitoring. TensorRT sits between training and deployment. NVIDIA says it optimizes models trained on all major frameworks, calibrates them for lower precision with high accuracy, and deploys them from data centers to edge devices.
What NVIDIA says (4)
“It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”
“TensorRT supports mixed precision”
“TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy, and deploy them to hyperscale data centers, workstations, laptops, and edge devices.”
“throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”
TensorRT is NVIDIA's SDK for optimizing deep learning inference on NVIDIA GPUs. Quantization stores numbers with fewer bits. Fusion merges several operations into one. Kernel tuning picks the fastest GPU code for each step. NVIDIA says TensorRT uses all three to optimize inference.
What NVIDIA says (2)
“is an SDK for optimizing deep learning inference on NVIDIA GPUs.”
“optimizes inference using quantization, layer and tensor fusion, and kernel tuning techniques.”
MLOps (machine learning operations) applies DevOps ideas to machine learning. NVIDIA defines it as practices and principles that streamline the development, deployment and maintenance of ML models in production. It covers the whole life cycle, from exploratory analysis and data preparation to deployment, monitoring and ongoing optimization.
What NVIDIA says (2)
“short for machine learning operations, is a set of practices and principles that aims to streamline the development, deployment, and maintenance of machine learning (ML) models in production environments.”
“MLOps is an extension of the existing discipline of DevOps, the modern practice of efficiently writing, deploying, and running enterprise applications.”
Training from scratch needs a lot of representative data and compute. NVIDIA says a pretrained model is trained on large datasets and can be used as is or fine-tuned to fit an application's needs. That reuses earlier work instead of repeating it.
What NVIDIA says (3)
“You can use the pretrained models for inference or fine-tune them with transfer learning.”
“It can be used as is or further fine-tuned to fit an application’s specific needs.”
“That model needs a lot of representative data to learn from.”
Key terms: Pretrained model MLOps
Try it: Training vs Inference Serving NVIDIA AI Software Stack
1.8 GPU vs CPU architecture
Why GPUs, with many cores running in parallel, suit AI better than CPUs for the heavy math.
Key points
NVIDIA's CUDA guide says a CPU is designed to run a serial sequence of operations as fast as possible. A GPU is designed to run thousands of threads in parallel, trading lower single-thread performance for much greater total throughput. GPUs devote more transistors to data processing, while CPUs devote more to caching and flow control.
What NVIDIA says (2)
“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”
“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”
A core is a processing unit. A thread is one sequence of instructions. NVIDIA says a CPU has a few cores with lots of cache memory, handling a few software threads at a time. A GPU has hundreds of cores that handle thousands of threads simultaneously.
What NVIDIA says (5)
“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”
“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”
“Architecturally, the CPU is composed of just a few cores with lots of cache memory that can handle a few software threads at a time.”
“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”
“In contrast, a GPU is composed of hundreds of cores that can handle thousands of threads simultaneously.”
A tensor is a multi-dimensional array of numbers, the basic data type of deep learning. NVIDIA says deep learning is accelerated by dedicated Tensor Cores in its GPUs. Its mixed-precision guide notes math runs much faster in reduced precision on GPUs with Tensor Core support.
What NVIDIA says (2)
“Tensor Cores were introduced in the NVIDIA Volta™ GPU architecture to accelerate matrix multiply and accumulate operations for machine learning and scientific applications.”
“Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”
NVIDIA defines accelerated computing as using specialized hardware to dramatically speed up work with parallel processing. It offloads demanding work that can bog down CPUs, which typically run tasks one after another. The CPU still runs the parts of the program that are serial.
What NVIDIA says (4)
“A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”
“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”
“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”
“It offloads demanding work that can bog down CPUs, processors that typically execute tasks in serial fashion.”
Memory bandwidth is how fast data moves between memory and the processor. NVIDIA's CUDA guide says a GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.
What NVIDIA says (1)
“A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”
Key terms: Mixed precision Tensor Core Graphics processing unit Accelerated computing
Try it: AI, ML & DL Foundations