1.1 Serving LLMs at scale
How to judge whether an LLM deployment is fast, scalable and reliable, and what limits it.
Key points
NVIDIA's inference-optimization post states that the two main contributors to graphics processing unit (GPU) memory during large language model (LLM) inference are the model weights and the key-value (KV) cache. The KV cache grows with batch size and sequence length, which is why bigger batches eventually overflow memory.
What NVIDIA says (2)
“In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”
“Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”
Requests in a batch generate different numbers of tokens, so with static batching every request waits for the longest one to finish. NVIDIA names in-flight batching as one way to reduce this waste.
What NVIDIA says (2)
“As a result, all requests in the batch must wait until the longest request is finished, which can be exacerbated by a large variance in the generation lengths.”
“There are methods to mitigate this, such as in-flight batching.”
Model parallelism means splitting one model across several graphics processing units (GPUs). NVIDIA notes this reduces the per-device memory footprint of the weights. Tensor parallelism splits individual layers into blocks on different devices. Pipeline parallelism places groups of layers on separate devices.
What NVIDIA says (3)
“One way to reduce the per-device memory footprint of the model weights is to distribute the model over several GPUs.”
“Pipeline parallelism involves sharding the model (vertically) into chunks, where each chunk comprises a subset of layers that is executed on a separate device.”
“Tensor parallelism involves sharding (horizontally) individual layers of the model into smaller, independent blocks of computation that can be executed on different devices.”
Key terms
- KV cache: Stored attention keys and values for earlier tokens; with model weights it dominates inference memory.
- In-flight batching: A serving technique that avoids making a whole batch wait for its longest request.
- Tensor / pipeline parallelism: Splitting one model across GPUs, by layer pieces (tensor) or by groups of layers (pipeline).
Try it
Sample question
A large language model (LLM) service batches requests to raise throughput, but larger batches start failing with out-of-memory errors. According to NVIDIA's inference-optimization guidance, which two things dominate graphics processing unit (GPU) memory during LLM inference?
Show the answer
Answer: Model weights and the KV cache
NVIDIA's inference-optimization post states that the two main contributors to graphics processing unit (GPU) memory during large language model (LLM) inference are the model weights and the key-value (KV) cache. The KV cache grows with batch size and sequence length, which is why bigger batches eventually overflow memory.
What NVIDIA says (2)
“In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”
“Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”
Practice 1.1 (3 questions) Full Core Machine Learning and AI Knowledge guide