1.6 NVIDIA solutions and what they are for
Which NVIDIA product solves which problem: NIM, Triton, NeMo, AI Enterprise, DGX and HGX.
Key points
A microservice is a small, self-contained service that other applications call over an API (application programming interface). NVIDIA NIM provides prebuilt, optimized inference microservices for deploying AI models on NVIDIA-accelerated infrastructure. NIM containers expose industry-standard APIs and are tuned for latency and throughput.
What NVIDIA says (3)
“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”
“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”
“expose industry-standard APIs for simple integration into AI applications, development frameworks, and workflows and optimize response latency and throughput for each combination of foundation model and GPU.”
NVIDIA describes the NeMo Framework as a scalable, cloud-native generative AI framework. It is built for researchers and developers working on LLMs, multimodal models and speech AI.
What NVIDIA says (1)
“NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI”
Inference serving means running trained models behind an endpoint that applications can call. Triton is open-source inference serving software. It deploys models from frameworks such as TensorRT, PyTorch and ONNX. Its features include concurrent model execution and dynamic batching, which groups requests to use the GPU better.
What NVIDIA says (3)
“Triton Inference Server is an open source inference serving software that streamlines AI inferencing.”
“Triton Inference Server enables teams to deploy any AI model from multiple deep learning and machine learning frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL, and more.”
“Concurrent model execution Dynamic batching”
NVIDIA AI Enterprise is NVIDIA's cloud-native software platform for production AI. NVIDIA says it increases reliability with extended-lifetime production branches and enterprise support. A production branch is a software release line kept stable and patched for a long time.
What NVIDIA says (1)
“Increase reliability with extended-lifetime production branches and enterprise support.”
cuDF is the RAPIDS library for DataFrame operations on the GPU. NVIDIA offers plug-in accelerators for DataFrame libraries like pandas with no code changes required.
What NVIDIA says (4)
“is a GPU-accelerated library for tabular data processing.”
“a zero-code change accelerator, cudf.pandas , for existing pandas code.”
“A Python library providing a GPU engine for Polars”
“accelerators for popular DataFrame libraries and SQL engines, like Polars, pandas, and Apache Spark with no code changes required.”
NVIDIA calls DGX a complete AI solution with full-stack, intelligent software. HGX brings together NVIDIA GPUs, CPUs, NVLink, networking and optimized software. NVIDIA says HGX is available as a single baseboard with eight GPUs, which server makers build systems around.
What NVIDIA says (3)
“Unlock productivity with the NVIDIA DGX platform, a complete AI solution with full-stack, intelligent software in NVIDIA Mission Control .”
“This server variant consists of one GPU baseboard with eight NVIDIA H100 GPUs and four NVSwitches.”
“NVIDIA HGX is available in a single baseboard with eight NVIDIA Rubin, NVIDIA Blackwell, or NVIDIA Blackwell Ultra SXMs.”
Key terms
- NVIDIA NIM: Prebuilt, optimized inference microservices for deploying AI models on NVIDIA GPUs anywhere.
- Triton Inference Server: Open-source software that serves models from many frameworks for inference.
Try it
Sample question
You need to deploy a trained model as a scalable inference microservice with industry-standard APIs. Which NVIDIA offering fits best?
Show the answer
Answer: NVIDIA NIM
A microservice is a small, self-contained service that other applications call over an API (application programming interface). NVIDIA NIM provides prebuilt, optimized inference microservices for deploying AI models on NVIDIA-accelerated infrastructure. NIM containers expose industry-standard APIs and are tuned for latency and throughput.
What NVIDIA says (3)
“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”
“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”
“expose industry-standard APIs for simple integration into AI applications, development frameworks, and workflows and optimize response latency and throughput for each combination of foundation model and GPU.”
Practice 1.6 (6 questions) Full Essential AI Knowledge guide
← 1.5 AI use cases and industries · 1.7 The AI development and deployment life cycle →