3.2 RAG, chatbots and summarizers

NCA-GENM · Multimodal Data (15% of the exam) · Official objective: “Build LLM use cases such as retrieval-augmented generation (RAG), chatbots, and summarizers.”

Multimodal RAG designs, the offline and online parts of RAG, and hallucination.

Key points

  1. CLIP can encode both text and images into the same space. The rest of the text RAG setup can stay much the same. A multimodal LLM (MLLM) then answers using retrieved images. CLIP means Contrastive Language-Image Pre-training. RAG means retrieval-augmented generation. LLM means large language model. MLLM means multimodal large language model.

    What NVIDIA says (2)

    “In the case of images and text, you can use a model like CLIP to encode both text and images in the same vector space.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

    “For the generation pass, you then replace the large language model (LLM) with a multimodal LLM (MLLM) for all question and answering.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

  2. Each approach trades simplicity against fidelity. Separate stores need a rank-rerank step to merge results. RAG means retrieval-augmented generation. OCR means optical character recognition.

    What NVIDIA says (1)

    “Embed all modalities into the same vector space Ground all modalities into one primary modality Have separate stores for different modalities”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

  3. Ingestion loads, splits and embeds documents into a vector database. Retrieval and generation happen online when a query arrives. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

  4. Hallucination is a plausible but incorrect answer. RAG supplies real sources as context, which lowers that risk. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”

    — What Is Retrieval-Augmented Generation aka RAG

Key terms

Try it

Sample question

You build multimodal RAG by embedding images and text into one vector space. What changes from a text-only pipeline?

Show the answer

Answer: Swap the embedding model for one like CLIP, and the LLM for a multimodal LLM

CLIP can encode both text and images into the same space. The rest of the text RAG setup can stay much the same. A multimodal LLM (MLLM) then answers using retrieved images. CLIP means Contrastive Language-Image Pre-training. RAG means retrieval-augmented generation. LLM means large language model. MLLM means multimodal large language model.

What NVIDIA says (2)

“In the case of images and text, you can use a model like CLIP to encode both text and images in the same vector space.”

— An Easy Introduction to Multimodal Retrieval-Augmented Generation

“For the generation pass, you then replace the large language model (LLM) with a multimodal LLM (MLLM) for all question and answering.”

— An Easy Introduction to Multimodal Retrieval-Augmented Generation

Practice 3.2 (4 questions) Full Multimodal Data guide

← 3.1 Scalability, performance and reliability · 3.3 Python natural-language packages →