1.1 The NVIDIA software stack
The layers between the GPU and your AI application: driver, CUDA, libraries, containers and serving.
Key points
The CUDA toolkit is what developers build with. The driver is what runs on the server. They are versioned separately. NVIDIA's CUDA Compatibility guide defines supported ways to run newer CUDA software on existing driver installations, within documented limits. Forward compatibility uses the cuda-compat package so applications built with a newer toolkit can run on older base drivers, subject to platform and GPU support.
What NVIDIA says (2)
“CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”
“It also includes Forward Compatibility , which uses the cuda-compat-<major>-<minor> package to allow applications built with a newer toolkit to run on older base drivers across major release families, subject to platform and GPU support.”
A collective is a communication step that involves every GPU in a job. All-reduce is a collective that sums values, such as gradients, across all GPUs and gives every GPU the result. NCCL (NVIDIA Collective Communications Library) implements multi-GPU and multi-node communication primitives. It provides all-gather, all-reduce, broadcast, reduce, reduce-scatter, and point-to-point send and receive.
What NVIDIA says (3)
“is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”
“NCCL implements both collective communication and point-to-point send/receive primitives.”
“NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”
CUDA is NVIDIA's platform for GPU computing. CUDA-X is the collection of libraries built on it. cuDNN (CUDA Deep Neural Network library) is a GPU-accelerated library of primitives for deep neural networks. A primitive is a basic building block that frameworks such as PyTorch call under the hood. NVIDIA lists attention, convolution, matrix multiplication, normalization and pooling among the operations it tunes.
What NVIDIA says (5)
“The NVIDIA CUDA Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives for deep neural networks.”
“It provides highly tuned implementations of operations arising frequently in deep neural network (DNN) applications: Scaled dot-product attention Convolution, including cross-correlation Matrix multiplication Normalizations, softmax, and pooling”
“to enable any computational workload to use the throughput capability of GPUs independent of graphics APIs.”
“Libraries like cuBLAS, cuFFT, cuDNN, and CUTLASS are just a few examples of libraries that help developers avoid reimplementing well-established algorithms.”
“NVIDIA CUDA-X™, built on CUDA, is a collection of libraries that deliver dramatically higher performance across application domains, including AI and HPC.”
RAPIDS is NVIDIA's set of CUDA-X data science libraries. cuDF accelerates fundamental DataFrame operations, the table operations pandas users run. cuML is a GPU-accelerated machine learning library. NVIDIA says zero-code-change APIs accelerate popular tools like pandas and scikit-learn.
What NVIDIA says (4)
“is a GPU-accelerated library for tabular data processing.”
“NVIDIA cuML is a suite of fast, GPU-accelerated machine learning algorithms designed for data science and analytical tasks.”
“a zero-code change accelerator, cudf.pandas , for existing pandas code.”
“automatically accelerate existing code with zero code changes.”
A container packages an application with everything it needs to run, so it behaves the same everywhere. The NGC catalog gives access to GPU-accelerated software: performance-optimized containers, pretrained AI models and industry-specific SDKs (software development kits). NVIDIA says these can be deployed on premises, in the cloud or at the edge.
What NVIDIA says (2)
“The NGC Catalog consists of containers, pretrained models, Helm charts for Kubernetes deployments, and industry-specific AI toolkits with software development kits (SDKs).”
“The NGC catalog provides access to GPU-accelerated software that speeds up end-to-end workflows with performance-optimized containers, pretrained AI models, and industry-specific SDKs that can be deployed on premises, in the cloud, or at the edge.”
Containers normally cannot see a server's GPUs. The NVIDIA Container Toolkit provides the runtime pieces that expose GPUs to containers. NVIDIA lists components such as the NVIDIA Container Runtime and the nvidia-ctk command-line tool. The GPU Operator installs it, together with the driver, on Kubernetes nodes.
What NVIDIA says (2)
“It currently includes: The NVIDIA Container Runtime ( nvidia-container-runtime ) The NVIDIA Container Toolkit CLI ( nvidia-ctk )”
“These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the NVIDIA Container Toolkit”
Key terms
- CUDA: NVIDIA's parallel computing platform; CUDA-X libraries built on it accelerate AI and HPC.
- cuDNN: A GPU library of tuned building blocks, such as convolutions, for deep neural networks.
- NCCL: The library that moves data between GPUs inside and across nodes, with operations such as all-reduce.
- NGC catalog: NVIDIA's catalog of GPU-optimized containers, pretrained models and SDKs.
- Container: A package of an application together with its software dependencies (libraries, configuration, tools) that runs the same on any host.
Try it
Sample question
A team wants to run an application built with a newer CUDA toolkit on servers whose GPU driver is older. What does NVIDIA's CUDA Compatibility guidance offer?
Show the answer
Answer: Supported ways to run newer CUDA software on existing drivers within documented limits, including forward compatibility through the cuda-compat package.
The CUDA toolkit is what developers build with. The driver is what runs on the server. They are versioned separately. NVIDIA's CUDA Compatibility guide defines supported ways to run newer CUDA software on existing driver installations, within documented limits. Forward compatibility uses the cuda-compat package so applications built with a newer toolkit can run on older base drivers, subject to platform and GPU support.
What NVIDIA says (2)
“CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”
“It also includes Forward Compatibility , which uses the cuda-compat-<major>-<minor> package to allow applications built with a newer toolkit to run on older base drivers across major release families, subject to platform and GPU support.”
Practice 1.1 (6 questions) Full Essential AI Knowledge guide