1.1 The NVIDIA software stack

NCA-AIIO · Essential AI Knowledge (38% of the exam) · Official objective: “Describe the NVIDIA software stack used in an AI environment.”

The layers between the GPU and your AI application: driver, CUDA, libraries, containers and serving.

Key points

  1. The CUDA toolkit is what developers build with. The driver is what runs on the server. They are versioned separately. NVIDIA's CUDA Compatibility guide defines supported ways to run newer CUDA software on existing driver installations, within documented limits. Forward compatibility uses the cuda-compat package so applications built with a newer toolkit can run on older base drivers, subject to platform and GPU support.

    What NVIDIA says (2)

    “CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”

    — NVIDIA CUDA Compatibility

    “It also includes Forward Compatibility , which uses the cuda-compat-<major>-<minor> package to allow applications built with a newer toolkit to run on older base drivers across major release families, subject to platform and GPU support.”

    — NVIDIA CUDA Compatibility

  2. A collective is a communication step that involves every GPU in a job. All-reduce is a collective that sums values, such as gradients, across all GPUs and gives every GPU the result. NCCL (NVIDIA Collective Communications Library) implements multi-GPU and multi-node communication primitives. It provides all-gather, all-reduce, broadcast, reduce, reduce-scatter, and point-to-point send and receive.

    What NVIDIA says (3)

    “is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”

    — Overview of NCCL — NCCL 2.32.3 documentation

    “NCCL implements both collective communication and point-to-point send/receive primitives.”

    — Overview of NCCL — NCCL 2.32.3 documentation

    “NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”

    — Overview of NCCL — NCCL 2.32.3 documentation

  3. CUDA is NVIDIA's platform for GPU computing. CUDA-X is the collection of libraries built on it. cuDNN (CUDA Deep Neural Network library) is a GPU-accelerated library of primitives for deep neural networks. A primitive is a basic building block that frameworks such as PyTorch call under the hood. NVIDIA lists attention, convolution, matrix multiplication, normalization and pooling among the operations it tunes.

    What NVIDIA says (5)

    “The NVIDIA CUDA Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives for deep neural networks.”

    — NVIDIA cuDNN Documentation

    “It provides highly tuned implementations of operations arising frequently in deep neural network (DNN) applications: Scaled dot-product attention Convolution, including cross-correlation Matrix multiplication Normalizations, softmax, and pooling”

    — NVIDIA cuDNN Documentation

    “to enable any computational workload to use the throughput capability of GPUs independent of graphics APIs.”

    — CUDA Programming Guide: Introduction

    “Libraries like cuBLAS, cuFFT, cuDNN, and CUTLASS are just a few examples of libraries that help developers avoid reimplementing well-established algorithms.”

    — CUDA Programming Guide: Introduction

    “NVIDIA CUDA-X™, built on CUDA, is a collection of libraries that deliver dramatically higher performance across application domains, including AI and HPC.”

    — CUDA Platform for Accelerated Computing

  4. RAPIDS is NVIDIA's set of CUDA-X data science libraries. cuDF accelerates fundamental DataFrame operations, the table operations pandas users run. cuML is a GPU-accelerated machine learning library. NVIDIA says zero-code-change APIs accelerate popular tools like pandas and scikit-learn.

    What NVIDIA says (4)

    “is a GPU-accelerated library for tabular data processing.”

    — NVIDIA cuDF Documentation

    “NVIDIA cuML is a suite of fast, GPU-accelerated machine learning algorithms designed for data science and analytical tasks.”

    — NVIDIA cuML Documentation

    “a zero-code change accelerator, cudf.pandas , for existing pandas code.”

    — NVIDIA cuDF Documentation

    “automatically accelerate existing code with zero code changes.”

    — NVIDIA cuML Documentation

  5. A container packages an application with everything it needs to run, so it behaves the same everywhere. The NGC catalog gives access to GPU-accelerated software: performance-optimized containers, pretrained AI models and industry-specific SDKs (software development kits). NVIDIA says these can be deployed on premises, in the cloud or at the edge.

    What NVIDIA says (2)

    “The NGC Catalog consists of containers, pretrained models, Helm charts for Kubernetes deployments, and industry-specific AI toolkits with software development kits (SDKs).”

    — NGC Catalog User Guide

    “The NGC catalog provides access to GPU-accelerated software that speeds up end-to-end workflows with performance-optimized containers, pretrained AI models, and industry-specific SDKs that can be deployed on premises, in the cloud, or at the edge.”

    — NVIDIA NGC Catalog

  6. Containers normally cannot see a server's GPUs. The NVIDIA Container Toolkit provides the runtime pieces that expose GPUs to containers. NVIDIA lists components such as the NVIDIA Container Runtime and the nvidia-ctk command-line tool. The GPU Operator installs it, together with the driver, on Kubernetes nodes.

    What NVIDIA says (2)

    “It currently includes: The NVIDIA Container Runtime ( nvidia-container-runtime ) The NVIDIA Container Toolkit CLI ( nvidia-ctk )”

    — NVIDIA Container Toolkit

    “These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the NVIDIA Container Toolkit”

    — About the NVIDIA GPU Operator

Key terms

Try it

Sample question

A team wants to run an application built with a newer CUDA toolkit on servers whose GPU driver is older. What does NVIDIA's CUDA Compatibility guidance offer?

Show the answer

Answer: Supported ways to run newer CUDA software on existing drivers within documented limits, including forward compatibility through the cuda-compat package.

The CUDA toolkit is what developers build with. The driver is what runs on the server. They are versioned separately. NVIDIA's CUDA Compatibility guide defines supported ways to run newer CUDA software on existing driver installations, within documented limits. Forward compatibility uses the cuda-compat package so applications built with a newer toolkit can run on older base drivers, subject to platform and GPU support.

What NVIDIA says (2)

“CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”

— NVIDIA CUDA Compatibility

“It also includes Forward Compatibility , which uses the cuda-compat-<major>-<minor> package to allow applications built with a newer toolkit to run on older base drivers across major release families, subject to platform and GPU support.”

— NVIDIA CUDA Compatibility

Practice 1.1 (6 questions) Full Essential AI Knowledge guide

1.2 Training vs inference →