1.1 Serving LLMs at scale

NCA-GENL · Core Machine Learning and AI Knowledge (30% of the exam) · Official objective: “Assist in deployment and evaluation of model scalability, performance, and reliability under the supervision of senior team members.”

How to judge whether an LLM deployment is fast, scalable and reliable, and what limits it.

Key points

  1. NVIDIA's inference-optimization post states that the two main contributors to graphics processing unit (GPU) memory during large language model (LLM) inference are the model weights and the key-value (KV) cache. The KV cache grows with batch size and sequence length, which is why bigger batches eventually overflow memory.

    What NVIDIA says (2)

    “In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”

    — Mastering LLM Techniques: Inference Optimization

    “Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”

    — Mastering LLM Techniques: Inference Optimization

  2. Requests in a batch generate different numbers of tokens, so with static batching every request waits for the longest one to finish. NVIDIA names in-flight batching as one way to reduce this waste.

    What NVIDIA says (2)

    “As a result, all requests in the batch must wait until the longest request is finished, which can be exacerbated by a large variance in the generation lengths.”

    — Mastering LLM Techniques: Inference Optimization

    “There are methods to mitigate this, such as in-flight batching.”

    — Mastering LLM Techniques: Inference Optimization

  3. Model parallelism means splitting one model across several graphics processing units (GPUs). NVIDIA notes this reduces the per-device memory footprint of the weights. Tensor parallelism splits individual layers into blocks on different devices. Pipeline parallelism places groups of layers on separate devices.

    What NVIDIA says (3)

    “One way to reduce the per-device memory footprint of the model weights is to distribute the model over several GPUs.”

    — Mastering LLM Techniques: Inference Optimization

    “Pipeline parallelism involves sharding the model (vertically) into chunks, where each chunk comprises a subset of layers that is executed on a separate device.”

    — Mastering LLM Techniques: Inference Optimization

    “Tensor parallelism involves sharding (horizontally) individual layers of the model into smaller, independent blocks of computation that can be executed on different devices.”

    — Mastering LLM Techniques: Inference Optimization

Key terms

Try it

Sample question

A large language model (LLM) service batches requests to raise throughput, but larger batches start failing with out-of-memory errors. According to NVIDIA's inference-optimization guidance, which two things dominate graphics processing unit (GPU) memory during LLM inference?

Show the answer

Answer: Model weights and the KV cache

NVIDIA's inference-optimization post states that the two main contributors to graphics processing unit (GPU) memory during large language model (LLM) inference are the model weights and the key-value (KV) cache. The KV cache grows with batch size and sequence length, which is why bigger batches eventually overflow memory.

What NVIDIA says (2)

“In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”

— Mastering LLM Techniques: Inference Optimization

“Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”

— Mastering LLM Techniques: Inference Optimization

Practice 1.1 (3 questions) Full Core Machine Learning and AI Knowledge guide

1.2 Finding insights in large datasets →