2.2 Building LLM features in software
Agents, reranking and context limits when you build LLM features.
Key points
A large language model (LLM)-powered agent uses an LLM to reason through a problem, create a plan and carry it out with tools. NVIDIA's example question needs planning, memory and several tools, which a simple lookup cannot provide.
What NVIDIA says (2)
“they can be described as a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”
“This inquiry requires planning, tailored focus, memory, using different tools”
Reranking is a second stage after fast retrieval such as Best Matching 25, a keyword-ranking function (BM25) or vector search. The candidates go to a large language model (LLM) that judges how relevant each one is to the query, so the best passages reach the generator.
What NVIDIA says (3)
“Initially, a set of candidate documents or passages is retrieved using traditional information retrieval methods like BM25 or vector similarity search.”
“Re-ranking is typically used as a second stage after an initial fast retrieval step”
“These candidates are then fed into an LLM that analyzes the semantic relevance between the query and each document.”
The context window is the most text the large language model (LLM) can take in at once. NVIDIA notes the whole prompt, retrieved chunks plus the query, must fit in it, so chunk sizes should not be too big.
What NVIDIA says (2)
“The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”
“specify chunk sizes too big”
NVIDIA's LLMOps post says that once a customized model is used alone or in a chain of models and application programming interfaces (APIs), you must test the complete AI system for accuracy, speed and vulnerabilities, and add guardrails.
What NVIDIA says (2)
“At this point, it is crucial to test the complete AI system for accuracy, speed”
“it is crucial to test the complete AI system for accuracy, speed, and vulnerabilities, and add guardrails”
Key terms
- Large language model: A deep-learning model trained on huge amounts of text that can understand and generate language.
- Retrieval-augmented generation: A pattern where relevant documents are retrieved at query time and given to the LLM as context for its answer.
- Reranking: A second retrieval stage that re-scores candidate passages for relevance to the query.
- Context window: The maximum amount of text (prompt plus retrieved context) an LLM can take in at once.
- LLM agent: A system that uses an LLM to reason, plan and act with tools.
Sample question
A question needs planning, memory and several tools (for example search plus a calculator) to answer. Which design fits?
Show the answer
Answer: An LLM-powered agent that reasons, plans and executes the plan with a set of tools
A large language model (LLM)-powered agent uses an LLM to reason through a problem, create a plan and carry it out with tools. NVIDIA's example question needs planning, memory and several tools, which a simple lookup cannot provide.
What NVIDIA says (2)
“they can be described as a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”
“This inquiry requires planning, tailored focus, memory, using different tools”
Practice 2.2 (4 questions) Full Software Development guide
← 2.1 Measuring LLM performance in software · 2.3 Python language packages in practice →