3.3 Testing LLM applications
How to test latency, behavior and safety of LLM apps.
Key points
Intertoken latency is the average time between consecutive generated tokens. NVIDIA notes tools differ in details; for example GenAI-Perf excludes time to first token from this average while LLMPerf includes it.
What NVIDIA says (2)
“For example, GenAI-Perf does not include TTFT in the average calculation (as opposed to LLMPerf, which does include the TTFT).”
“is the average time between the generation of consecutive tokens in a sequence. It is also known as time per output token (TPOT).”
Red teaming means assessing a system the way an attacker would, to find and reduce risks. NVIDIA's AI red team combines offensive-security professionals and data scientists to assess machine learning (ML) systems and help mitigate risks.
What NVIDIA says (2)
“Our AI red team is a cross-functional team made up of offensive security professionals and data scientists. We use our combined skills to assess our ML systems to identify and help mitigate any risks”
“mitigate any risks from the perspective of information security.”
Key terms
- Time to first token: How long until the first output token arrives.
- Intertoken latency: The average time between consecutive output tokens.
- Red teaming: Assessing an AI system from an attacker's point of view to find and reduce risks.
Sample question
What does intertoken latency measure in large language model (LLM) benchmarking?
Show the answer
Answer: The average time between consecutive output tokens
Intertoken latency is the average time between consecutive generated tokens. NVIDIA notes tools differ in details; for example GenAI-Perf excludes time to first token from this average while LLMPerf includes it.
What NVIDIA says (2)
“For example, GenAI-Perf does not include TTFT in the average calculation (as opposed to LLMPerf, which does include the TTFT).”
“is the average time between the generation of consecutive tokens in a sequence. It is also known as time per output token (TPOT).”
Practice 3.3 (2 questions) Full Experimentation guide
← 3.2 Preparing large datasets · 3.4 Evaluating technologies →