2.4 Picking the right components

NCA-GENL · Software Development (24% of the exam) · Official objective: “Identify system data, hardware, or software components required to meet user needs.”

Which NVIDIA software components fit which need.

Key points

  1. NIM (NVIDIA Inference Microservices) are containerized solutions with industry-standard application programming interfaces (APIs) and Helm charts to scale. They let developers focus on application logic instead of serving infrastructure.

    What NVIDIA says (2)

    “NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”

    — Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

    “developers to focus on working on their application logic rather than having to spend cycles on building and scaling out the infrastructure.”

    — Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

  2. NVIDIA Collective Communications Library (NCCL) provides inter-GPU communication (communication between graphics processing units, or GPUs) primitives that are topology-aware. Its AllReduce collective is heavily used in neural-network training.

    What NVIDIA says (2)

    “is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”

    — Overview of NCCL — NCCL 2.32.3 documentation

    “NCCL has found great application in Deep Learning Frameworks, where the AllReduce collective is heavily used for neural network training.”

    — Overview of NCCL — NCCL 2.32.3 documentation

  3. NVIDIA NeMo Retriever is a collection of microservices with a single application programming interface (API) for indexing and querying user data. It covers extraction, embedding and reranking pipelines.

    What NVIDIA says (2)

    “NVIDIA NeMo Retriever (NeMo Retriever) is a collection of microservices that present a single API for indexing and querying of user data.”

    — NVIDIA NeMo Retriever

    “for building and scaling multimodal data extraction, embedding, and reranking pipelines”

    — NVIDIA NeMo Retriever

  4. NVIDIA NeMo Framework is a scalable, cloud-native generative AI framework for large language models (LLMs), multimodal and speech models. It lets users create, customize and deploy models from existing code and pre-trained checkpoints.

    What NVIDIA says (2)

    “NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and”

    — Overview — NVIDIA NeMo Framework User Guide

    “It enables users to efficiently create, customize, and deploy new generative AI models by leveraging existing code and pre-trained model checkpoints.”

    — Overview — NVIDIA NeMo Framework User Guide

  5. FP16 (half precision) uses 16 bits instead of 32 bits for FP32 (single precision). NVIDIA's mixed-precision guide says lowering memory enables larger models or larger mini-batches.

    What NVIDIA says (1)

    “Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32). Lowering the required memory enables training of larger models or training with larger mini-batches.”

    — Train With Mixed Precision

Key terms

Sample question

A team wants to deploy a model behind a standard application programming interface (API) without building and scaling its own serving infrastructure. Which NVIDIA component is designed for this?

Show the answer

Answer: NVIDIA NIM inference microservices

NIM (NVIDIA Inference Microservices) are containerized solutions with industry-standard application programming interfaces (APIs) and Helm charts to scale. They let developers focus on application logic instead of serving infrastructure.

What NVIDIA says (2)

“NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”

— Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

“developers to focus on working on their application logic rather than having to spend cycles on building and scaling out the infrastructure.”

— Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

Practice 2.4 (5 questions) Full Software Development guide

← 2.3 Python language packages in practice · 2.5 Monitoring data, experiments and processes →