1.6 NVIDIA solutions and what they are for

NCA-AIIO · Essential AI Knowledge (38% of the exam) · Official objective: “Explain the purpose and use case of various NVIDIA solutions.”

Which NVIDIA product solves which problem: NIM, Triton, NeMo, AI Enterprise, DGX and HGX.

Key points

  1. A microservice is a small, self-contained service that other applications call over an API (application programming interface). NVIDIA NIM provides prebuilt, optimized inference microservices for deploying AI models on NVIDIA-accelerated infrastructure. NIM containers expose industry-standard APIs and are tuned for latency and throughput.

    What NVIDIA says (3)

    “NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”

    — NVIDIA NIM Documentation

    “NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”

    — NVIDIA NIM Microservices

    “expose industry-standard APIs for simple integration into AI applications, development frameworks, and workflows and optimize response latency and throughput for each combination of foundation model and GPU.”

    — NVIDIA NIM for Developers

  2. NVIDIA describes the NeMo Framework as a scalable, cloud-native generative AI framework. It is built for researchers and developers working on LLMs, multimodal models and speech AI.

    What NVIDIA says (1)

    “NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI”

    — NVIDIA NeMo Framework

  3. Inference serving means running trained models behind an endpoint that applications can call. Triton is open-source inference serving software. It deploys models from frameworks such as TensorRT, PyTorch and ONNX. Its features include concurrent model execution and dynamic batching, which groups requests to use the GPU better.

    What NVIDIA says (3)

    “Triton Inference Server is an open source inference serving software that streamlines AI inferencing.”

    — NVIDIA Triton Inference Server

    “Triton Inference Server enables teams to deploy any AI model from multiple deep learning and machine learning frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL, and more.”

    — NVIDIA Triton Inference Server

    “Concurrent model execution Dynamic batching”

    — NVIDIA Triton Inference Server

  4. NVIDIA AI Enterprise is NVIDIA's cloud-native software platform for production AI. NVIDIA says it increases reliability with extended-lifetime production branches and enterprise support. A production branch is a software release line kept stable and patched for a long time.

    What NVIDIA says (1)

    “Increase reliability with extended-lifetime production branches and enterprise support.”

    — NVIDIA AI Enterprise

  5. cuDF is the RAPIDS library for DataFrame operations on the GPU. NVIDIA offers plug-in accelerators for DataFrame libraries like pandas with no code changes required.

    What NVIDIA says (4)

    “is a GPU-accelerated library for tabular data processing.”

    — NVIDIA cuDF Documentation

    “a zero-code change accelerator, cudf.pandas , for existing pandas code.”

    — NVIDIA cuDF Documentation

    “A Python library providing a GPU engine for Polars”

    — NVIDIA cuDF Documentation

    “accelerators for popular DataFrame libraries and SQL engines, like Polars, pandas, and Apache Spark with no code changes required.”

    — CUDA-X Data Science Libraries (RAPIDS)

  6. NVIDIA calls DGX a complete AI solution with full-stack, intelligent software. HGX brings together NVIDIA GPUs, CPUs, NVLink, networking and optimized software. NVIDIA says HGX is available as a single baseboard with eight GPUs, which server makers build systems around.

    What NVIDIA says (3)

    “Unlock productivity with the NVIDIA DGX platform, a complete AI solution with full-stack, intelligent software in NVIDIA Mission Control .”

    — NVIDIA DGX Platform

    “This server variant consists of one GPU baseboard with eight NVIDIA H100 GPUs and four NVSwitches.”

    — NVIDIA Fabric Manager User Guide

    “NVIDIA HGX is available in a single baseboard with eight NVIDIA Rubin, NVIDIA Blackwell, or NVIDIA Blackwell Ultra SXMs.”

    — NVIDIA HGX Platform

Key terms

Try it

Sample question

You need to deploy a trained model as a scalable inference microservice with industry-standard APIs. Which NVIDIA offering fits best?

Show the answer

Answer: NVIDIA NIM

A microservice is a small, self-contained service that other applications call over an API (application programming interface). NVIDIA NIM provides prebuilt, optimized inference microservices for deploying AI models on NVIDIA-accelerated infrastructure. NIM containers expose industry-standard APIs and are tuned for latency and throughput.

What NVIDIA says (3)

“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”

— NVIDIA NIM Documentation

“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”

— NVIDIA NIM Microservices

“expose industry-standard APIs for simple integration into AI applications, development frameworks, and workflows and optimize response latency and throughput for each combination of foundation model and GPU.”

— NVIDIA NIM for Developers

Practice 1.6 (6 questions) Full Essential AI Knowledge guide

← 1.5 AI use cases and industries · 1.7 The AI development and deployment life cycle →