2.1 Measuring LLM performance in software

NCA-GENL · Software Development (24% of the exam) · Official objective: “Assist in the deployment and evaluations of model scalability, performance, and reliability under the supervision of a senior team member.”

Latency, throughput and testing approaches for LLM services.

Key points

  1. Time to first token (TTFT) is how long a user waits before the first output token appears. NVIDIA says TTFT generally includes request queuing, prefill and network latency. A longer prompt gives a larger TTFT because attention needs the whole input sequence.

    What NVIDIA says (2)

    “TTFT generally includes both request queuing time, prefill time, and network latency. The longer the prompt, the larger the TTFT.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “This is because the attention mechanism requires the whole input sequence to compute”

    — LLM Inference Benchmarking: Fundamental Concepts

  2. NVIDIA describes load testing as simulating many concurrent requests to assess handling of real-world traffic at scale. Performance benchmarking, for example with the NVIDIA GenAI-Perf tool, measures the model's own throughput, latency and token-level metrics. NVIDIA recommends combining both.

    What NVIDIA says (3)

    “Load testing focuses on simulating a large number of concurrent requests to a model to assess its ability to handle real-world traffic at scale.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “In contrast, performance benchmarking, as demonstrated by the NVIDIA GenAI-Perf tool, is concerned with measuring the actual performance of the model itself, such as its throughput, latency, and token-level metrics.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “By combining both approaches, developers can gain a comprehensive understanding”

    — LLM Inference Benchmarking: Fundamental Concepts

  3. NVIDIA's benchmarking post says the cost of a large language model (LLM) deployment depends on how many queries it can process per second while staying responsive. So serving performance drives cost.

    What NVIDIA says (1)

    “The cost of an LLM application deployment depends on how many queries it can process per second while being responsive”

    — LLM Inference Benchmarking: Fundamental Concepts

Key terms

Try it

Sample question

When benchmarking a large language model (LLM) endpoint, what does time to first token (TTFT) measure, and why does it grow with longer prompts?

Show the answer

Answer: Time until the first output token arrives, including queuing, prefill and network latency; longer prompts need more prefill work

Time to first token (TTFT) is how long a user waits before the first output token appears. NVIDIA says TTFT generally includes request queuing, prefill and network latency. A longer prompt gives a larger TTFT because attention needs the whole input sequence.

What NVIDIA says (2)

“TTFT generally includes both request queuing time, prefill time, and network latency. The longer the prompt, the larger the TTFT.”

— LLM Inference Benchmarking: Fundamental Concepts

“This is because the attention mechanism requires the whole input sequence to compute”

— LLM Inference Benchmarking: Fundamental Concepts

Practice 2.1 (3 questions) Full Software Development guide

← 1.10 Traditional ML with Python packages · 2.2 Building LLM features in software →