1.3 Building LLM applications: RAG, chatbots, summarizers
How the common LLM application patterns fit together.
Key points
RAG (retrieval-augmented generation) retrieves relevant passages at query time and gives them to the large language model (LLM). NVIDIA notes it gives models sources they can cite, is faster and cheaper than retraining, and lets you hot-swap new sources.
What NVIDIA says (3)
“Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”
“That makes the method faster and less expensive than retraining a model with additional datasets. And it lets users hot-swap new sources on the fly.”
“when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
NVIDIA calls a very plausible but incorrect answer a hallucination. It says retrieval-augmented generation (RAG) reduces the chance of hallucination.
What NVIDIA says (1)
“It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”
NVIDIA's retrieval-augmented generation (RAG) 101 post explains that document ingestion happens offline, while retrieval and response generation happen when an online query comes in.
What NVIDIA says (1)
“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
In retrieval-augmented generation (RAG), the large language model (LLM) is the generative component: it writes the answer using the user query and the contextual information retrieved from the vector database.
What NVIDIA says (1)
“In the context of RAG, LLMs are used to generate fully formed responses based on the user query and contextual information retrieved from the vector DBs during user queries.”
Key terms
- Large language model: A deep-learning model trained on huge amounts of text that can understand and generate language.
- Retrieval-augmented generation: A pattern where relevant documents are retrieved at query time and given to the LLM as context for its answer.
- Hallucination: A plausible-sounding but incorrect answer from a model.
Try it
Sample question
A team wants an assistant that answers questions about internal documents that change every week, with citations users can check. Which approach fits best?
Show the answer
Answer: Retrieval-augmented generation: retrieve relevant passages at query time and give them to the LLM
RAG (retrieval-augmented generation) retrieves relevant passages at query time and gives them to the large language model (LLM). NVIDIA notes it gives models sources they can cite, is faster and cheaper than retraining, and lets you hot-swap new sources.
What NVIDIA says (3)
“Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”
“That makes the method faster and less expensive than retraining a model with additional datasets. And it lets users hot-swap new sources on the fly.”
“when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
Practice 1.3 (4 questions) Full Core Machine Learning and AI Knowledge guide
← 1.2 Finding insights in large datasets · 1.4 Preparing content for RAG →