3.7 Running experiments on models and pipelines
Designing experiments that tell you what to change.
Key points
NVIDIA's chunking research tested strategies across several datasets. Even within the same document category, the best strategy varied significantly. So test chunking on your own content.
What NVIDIA says (2)
“Our research evaluated different chunking strategies across multiple datasets to establish guidelines for selecting the optimal approach based on your specific content and use case.”
“Even within the same document category, optimal chunking strategies varied significantly.”
Standardized benchmarks give consistent datasets and metrics. NVIDIA warns that academic benchmarks can become saturated quickly as large language models (LLMs) improve, and tailored benchmarks for specific domains are often missing.
What NVIDIA says (3)
“Note that as LLMs develop quickly, academic benchmarks can become saturated quick”
“The lack of tailored benchmarks for specific domains limits the relevance and depth of evaluations”
“Standardized benchmarks provide consistent datasets and metrics to evaluate LLMs across a variety of tasks.”
Sample question
In NVIDIA's chunking experiments, what did they find about the best chunk size?
Show the answer
Answer: The best strategy varied by dataset and query type, so it should be tested on your own content
NVIDIA's chunking research tested strategies across several datasets. Even within the same document category, the best strategy varied significantly. So test chunking on your own content.
What NVIDIA says (2)
“Our research evaluated different chunking strategies across multiple datasets to establish guidelines for selecting the optimal approach based on your specific content and use case.”
“Even within the same document category, optimal chunking strategies varied significantly.”
Practice 3.7 (2 questions) Full Experimentation guide
← 3.6 Evaluating and benchmarking models · 4.1 Insights from large datasets →