1.4 Preparing content for RAG

NCA-GENL · Core Machine Learning and AI Knowledge (30% of the exam) · Official objective: “Curate and embed content datasets for RAGs.”

How documents are split, embedded and stored so retrieval works.

Key points

  1. NVIDIA notes that text splitting breaks long text into smaller segments, which is necessary for fitting the text into the embedding model; the retrieval step then works on these smaller chunks rather than whole documents.

    What NVIDIA says (2)

    “One transformation method is text-splitting, which breaks down long text into smaller segments. This is necessary for fitting the text into the embedding model,”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

    “The retrieval step typically examines smaller chunks of the original text rather than all documents.”

    — Enhancing RAG Pipelines with Re-Ranking

  2. NVIDIA's chunking study says poor chunking can lead to irrelevant or incomplete responses. A smart strategy improves retrieval precision and contextual coherence.

    What NVIDIA says (2)

    “When done poorly, chunking can lead to irrelevant or incomplete responses, frustrating users and undermining trust in the system.”

    — Finding the Best Chunking Strategy for Accurate AI Responses

    “a smart chunking strategy improves retrieval precision and contextual coherence”

    — Finding the Best Chunking Strategy for Accurate AI Responses

  3. NVIDIA's vector-database glossary: ingested data is chunked, a vector is created to represent each chunk, and chunks and vectors are stored together with optional metadata.

    What NVIDIA says (1)

    “When private enterprise data is ingested, it’s chunked, a vector is created to represent it, and the data chunks with their corresponding vectors are stored in a vector database along with optional metadata for later retrieval.”

    — What is a Vector Database and How Does it Work?

Key terms

Try it

Sample question

During retrieval-augmented generation (RAG) ingestion, why are long documents split into smaller chunks before embedding?

Show the answer

Answer: So the text fits the embedding model and can be retrieved as focused pieces

NVIDIA notes that text splitting breaks long text into smaller segments, which is necessary for fitting the text into the embedding model; the retrieval step then works on these smaller chunks rather than whole documents.

What NVIDIA says (2)

“One transformation method is text-splitting, which breaks down long text into smaller segments. This is necessary for fitting the text into the embedding model,”

— RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

“The retrieval step typically examines smaller chunks of the original text rather than all documents.”

— Enhancing RAG Pipelines with Re-Ranking

Practice 1.4 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 1.3 Building LLM applications: RAG, chatbots, summarizers · 1.5 Machine-learning fundamentals →