2.4 Picking the right components
Which NVIDIA software components fit which need.
Key points
NIM (NVIDIA Inference Microservices) are containerized solutions with industry-standard application programming interfaces (APIs) and Helm charts to scale. They let developers focus on application logic instead of serving infrastructure.
What NVIDIA says (2)
“NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”
“developers to focus on working on their application logic rather than having to spend cycles on building and scaling out the infrastructure.”
NVIDIA Collective Communications Library (NCCL) provides inter-GPU communication (communication between graphics processing units, or GPUs) primitives that are topology-aware. Its AllReduce collective is heavily used in neural-network training.
What NVIDIA says (2)
“is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”
“NCCL has found great application in Deep Learning Frameworks, where the AllReduce collective is heavily used for neural network training.”
NVIDIA NeMo Retriever is a collection of microservices with a single application programming interface (API) for indexing and querying user data. It covers extraction, embedding and reranking pipelines.
What NVIDIA says (2)
“NVIDIA NeMo Retriever (NeMo Retriever) is a collection of microservices that present a single API for indexing and querying of user data.”
“for building and scaling multimodal data extraction, embedding, and reranking pipelines”
NVIDIA NeMo Framework is a scalable, cloud-native generative AI framework for large language models (LLMs), multimodal and speech models. It lets users create, customize and deploy models from existing code and pre-trained checkpoints.
What NVIDIA says (2)
“NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and”
“It enables users to efficiently create, customize, and deploy new generative AI models by leveraging existing code and pre-trained model checkpoints.”
FP16 (half precision) uses 16 bits instead of 32 bits for FP32 (single precision). NVIDIA's mixed-precision guide says lowering memory enables larger models or larger mini-batches.
What NVIDIA says (1)
“Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32). Lowering the required memory enables training of larger models or training with larger mini-batches.”
Key terms
- FP16 / FP32: 16-bit and 32-bit floating-point number formats.
- NVIDIA NIM: Containerized model-serving microservices with industry-standard APIs.
- NVIDIA NeMo Framework: NVIDIA's framework to create, customize and deploy generative AI models.
- NeMo Retriever: NVIDIA microservices for indexing and querying data, with embedding and reranking.
- NCCL: A library for fast communication between GPUs, such as AllReduce.
Sample question
A team wants to deploy a model behind a standard application programming interface (API) without building and scaling its own serving infrastructure. Which NVIDIA component is designed for this?
Show the answer
Answer: NVIDIA NIM inference microservices
NIM (NVIDIA Inference Microservices) are containerized solutions with industry-standard application programming interfaces (APIs) and Helm charts to scale. They let developers focus on application logic instead of serving infrastructure.
What NVIDIA says (2)
“NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”
“developers to focus on working on their application logic rather than having to spend cycles on building and scaling out the infrastructure.”
Practice 2.4 (5 questions) Full Software Development guide
← 2.3 Python language packages in practice · 2.5 Monitoring data, experiments and processes →