2.1 Measuring LLM performance in software
Latency, throughput and testing approaches for LLM services.
Key points
Time to first token (TTFT) is how long a user waits before the first output token appears. NVIDIA says TTFT generally includes request queuing, prefill and network latency. A longer prompt gives a larger TTFT because attention needs the whole input sequence.
What NVIDIA says (2)
“TTFT generally includes both request queuing time, prefill time, and network latency. The longer the prompt, the larger the TTFT.”
“This is because the attention mechanism requires the whole input sequence to compute”
NVIDIA describes load testing as simulating many concurrent requests to assess handling of real-world traffic at scale. Performance benchmarking, for example with the NVIDIA GenAI-Perf tool, measures the model's own throughput, latency and token-level metrics. NVIDIA recommends combining both.
What NVIDIA says (3)
“Load testing focuses on simulating a large number of concurrent requests to a model to assess its ability to handle real-world traffic at scale.”
“In contrast, performance benchmarking, as demonstrated by the NVIDIA GenAI-Perf tool, is concerned with measuring the actual performance of the model itself, such as its throughput, latency, and token-level metrics.”
“By combining both approaches, developers can gain a comprehensive understanding”
NVIDIA's benchmarking post says the cost of a large language model (LLM) deployment depends on how many queries it can process per second while staying responsive. So serving performance drives cost.
What NVIDIA says (1)
“The cost of an LLM application deployment depends on how many queries it can process per second while being responsive”
Key terms
- Time to first token: How long until the first output token arrives.
Try it
Sample question
When benchmarking a large language model (LLM) endpoint, what does time to first token (TTFT) measure, and why does it grow with longer prompts?
Show the answer
Answer: Time until the first output token arrives, including queuing, prefill and network latency; longer prompts need more prefill work
Time to first token (TTFT) is how long a user waits before the first output token appears. NVIDIA says TTFT generally includes request queuing, prefill and network latency. A longer prompt gives a larger TTFT because attention needs the whole input sequence.
What NVIDIA says (2)
“TTFT generally includes both request queuing time, prefill time, and network latency. The longer the prompt, the larger the TTFT.”
“This is because the attention mechanism requires the whole input sequence to compute”
Practice 2.1 (3 questions) Full Software Development guide
← 1.10 Traditional ML with Python packages · 2.2 Building LLM features in software →