3.2 RAG, chatbots and summarizers
Multimodal RAG designs, the offline and online parts of RAG, and hallucination.
Key points
CLIP can encode both text and images into the same space. The rest of the text RAG setup can stay much the same. A multimodal LLM (MLLM) then answers using retrieved images. CLIP means Contrastive Language-Image Pre-training. RAG means retrieval-augmented generation. LLM means large language model. MLLM means multimodal large language model.
What NVIDIA says (2)
“In the case of images and text, you can use a model like CLIP to encode both text and images in the same vector space.”
“For the generation pass, you then replace the large language model (LLM) with a multimodal LLM (MLLM) for all question and answering.”
Each approach trades simplicity against fidelity. Separate stores need a rank-rerank step to merge results. RAG means retrieval-augmented generation. OCR means optical character recognition.
What NVIDIA says (1)
“Embed all modalities into the same vector space Ground all modalities into one primary modality Have separate stores for different modalities”
Ingestion loads, splits and embeds documents into a vector database. Retrieval and generation happen online when a query arrives. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
Hallucination is a plausible but incorrect answer. RAG supplies real sources as context, which lowers that risk. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”
Key terms
- Modality: One type of data, such as text, image, audio or video.
- Retrieval-augmented generation: A method that retrieves relevant data at query time and gives it to the model as context.
- Hallucination: A plausible but incorrect answer from a generative model.
Try it
Sample question
You build multimodal RAG by embedding images and text into one vector space. What changes from a text-only pipeline?
Show the answer
Answer: Swap the embedding model for one like CLIP, and the LLM for a multimodal LLM
CLIP can encode both text and images into the same space. The rest of the text RAG setup can stay much the same. A multimodal LLM (MLLM) then answers using retrieved images. CLIP means Contrastive Language-Image Pre-training. RAG means retrieval-augmented generation. LLM means large language model. MLLM means multimodal large language model.
What NVIDIA says (2)
“In the case of images and text, you can use a model like CLIP to encode both text and images in the same vector space.”
“For the generation pass, you then replace the large language model (LLM) with a multimodal LLM (MLLM) for all question and answering.”
Practice 3.2 (4 questions) Full Multimodal Data guide
← 3.1 Scalability, performance and reliability · 3.3 Python natural-language packages →