3.3 Testing LLM applications

NCA-GENL · Experimentation (22% of the exam) · Official objective: “Assist in the design and conduct of hardware or software tests for LLM applications.”

How to test latency, behavior and safety of LLM apps.

Key points

  1. Intertoken latency is the average time between consecutive generated tokens. NVIDIA notes tools differ in details; for example GenAI-Perf excludes time to first token from this average while LLMPerf includes it.

    What NVIDIA says (2)

    “For example, GenAI-Perf does not include TTFT in the average calculation (as opposed to LLMPerf, which does include the TTFT).”

    — LLM Inference Benchmarking: Fundamental Concepts

    “is the average time between the generation of consecutive tokens in a sequence. It is also known as time per output token (TPOT).”

    — LLM Inference Benchmarking: Fundamental Concepts

  2. Red teaming means assessing a system the way an attacker would, to find and reduce risks. NVIDIA's AI red team combines offensive-security professionals and data scientists to assess machine learning (ML) systems and help mitigate risks.

    What NVIDIA says (2)

    “Our AI red team is a cross-functional team made up of offensive security professionals and data scientists. We use our combined skills to assess our ML systems to identify and help mitigate any risks”

    — NVIDIA AI Red Team: An Introduction

    “mitigate any risks from the perspective of information security.”

    — NVIDIA AI Red Team: An Introduction

Key terms

Sample question

What does intertoken latency measure in large language model (LLM) benchmarking?

Show the answer

Answer: The average time between consecutive output tokens

Intertoken latency is the average time between consecutive generated tokens. NVIDIA notes tools differ in details; for example GenAI-Perf excludes time to first token from this average while LLMPerf includes it.

What NVIDIA says (2)

“For example, GenAI-Perf does not include TTFT in the average calculation (as opposed to LLMPerf, which does include the TTFT).”

— LLM Inference Benchmarking: Fundamental Concepts

“is the average time between the generation of consecutive tokens in a sequence. It is also known as time per output token (TPOT).”

— LLM Inference Benchmarking: Fundamental Concepts

Practice 3.3 (2 questions) Full Experimentation guide

← 3.2 Preparing large datasets · 3.4 Evaluating technologies →