2.2 Building LLM features in software

NCA-GENL · Software Development (24% of the exam) · Official objective: “Build LLM use cases such as RAGs, chatbots, and summarizers.”

Agents, reranking and context limits when you build LLM features.

Key points

  1. A large language model (LLM)-powered agent uses an LLM to reason through a problem, create a plan and carry it out with tools. NVIDIA's example question needs planning, memory and several tools, which a simple lookup cannot provide.

    What NVIDIA says (2)

    “they can be described as a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”

    — Introduction to LLM Agents

    “This inquiry requires planning, tailored focus, memory, using different tools”

    — Introduction to LLM Agents

  2. Reranking is a second stage after fast retrieval such as Best Matching 25, a keyword-ranking function (BM25) or vector search. The candidates go to a large language model (LLM) that judges how relevant each one is to the query, so the best passages reach the generator.

    What NVIDIA says (3)

    “Initially, a set of candidate documents or passages is retrieved using traditional information retrieval methods like BM25 or vector similarity search.”

    — Enhancing RAG Pipelines with Re-Ranking

    “Re-ranking is typically used as a second stage after an initial fast retrieval step”

    — Enhancing RAG Pipelines with Re-Ranking

    “These candidates are then fed into an LLM that analyzes the semantic relevance between the query and each document.”

    — Enhancing RAG Pipelines with Re-Ranking

  3. The context window is the most text the large language model (LLM) can take in at once. NVIDIA notes the whole prompt, retrieved chunks plus the query, must fit in it, so chunk sizes should not be too big.

    What NVIDIA says (2)

    “The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”

    — Enhancing RAG Pipelines with Re-Ranking

    “specify chunk sizes too big”

    — Enhancing RAG Pipelines with Re-Ranking

  4. NVIDIA's LLMOps post says that once a customized model is used alone or in a chain of models and application programming interfaces (APIs), you must test the complete AI system for accuracy, speed and vulnerabilities, and add guardrails.

    What NVIDIA says (2)

    “At this point, it is crucial to test the complete AI system for accuracy, speed”

    — Mastering LLM Techniques: LLMOps

    “it is crucial to test the complete AI system for accuracy, speed, and vulnerabilities, and add guardrails”

    — Mastering LLM Techniques: LLMOps

Key terms

Sample question

A question needs planning, memory and several tools (for example search plus a calculator) to answer. Which design fits?

Show the answer

Answer: An LLM-powered agent that reasons, plans and executes the plan with a set of tools

A large language model (LLM)-powered agent uses an LLM to reason through a problem, create a plan and carry it out with tools. NVIDIA's example question needs planning, memory and several tools, which a simple lookup cannot provide.

What NVIDIA says (2)

“they can be described as a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”

— Introduction to LLM Agents

“This inquiry requires planning, tailored focus, memory, using different tools”

— Introduction to LLM Agents

Practice 2.2 (4 questions) Full Software Development guide

← 2.1 Measuring LLM performance in software · 2.3 Python language packages in practice →