1.4 Preparing content for RAG
How documents are split, embedded and stored so retrieval works.
Key points
NVIDIA notes that text splitting breaks long text into smaller segments, which is necessary for fitting the text into the embedding model; the retrieval step then works on these smaller chunks rather than whole documents.
What NVIDIA says (2)
“One transformation method is text-splitting, which breaks down long text into smaller segments. This is necessary for fitting the text into the embedding model,”
“The retrieval step typically examines smaller chunks of the original text rather than all documents.”
NVIDIA's chunking study says poor chunking can lead to irrelevant or incomplete responses. A smart strategy improves retrieval precision and contextual coherence.
What NVIDIA says (2)
“When done poorly, chunking can lead to irrelevant or incomplete responses, frustrating users and undermining trust in the system.”
“a smart chunking strategy improves retrieval precision and contextual coherence”
NVIDIA's vector-database glossary: ingested data is chunked, a vector is created to represent each chunk, and chunks and vectors are stored together with optional metadata.
What NVIDIA says (1)
“When private enterprise data is ingested, it’s chunked, a vector is created to represent it, and the data chunks with their corresponding vectors are stored in a vector database along with optional metadata for later retrieval.”
Key terms
- Vector database: A database that stores embeddings and finds the ones most similar to a query.
- Retrieval-augmented generation: A pattern where relevant documents are retrieved at query time and given to the LLM as context for its answer.
- Chunking: Breaking long documents into smaller pieces before embedding and retrieval.
Try it
Sample question
During retrieval-augmented generation (RAG) ingestion, why are long documents split into smaller chunks before embedding?
Show the answer
Answer: So the text fits the embedding model and can be retrieved as focused pieces
NVIDIA notes that text splitting breaks long text into smaller segments, which is necessary for fitting the text into the embedding model; the retrieval step then works on these smaller chunks rather than whole documents.
What NVIDIA says (2)
“One transformation method is text-splitting, which breaks down long text into smaller segments. This is necessary for fitting the text into the embedding model,”
“The retrieval step typically examines smaller chunks of the original text rather than all documents.”
Practice 1.4 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 1.3 Building LLM applications: RAG, chatbots, summarizers · 1.5 Machine-learning fundamentals →