1.2 Training vs inference
What training and inference each do, and why they need different resources.
Key points
Training is how a model learns. Data passes through the network, the error is propagated back through its layers, and the model adjusts its weights and tries again. A weight is one of the numbers inside the model that training tunes. Inference is when a trained model makes predictions on new data. NVIDIA notes inference cannot happen without training.
What NVIDIA says (3)
“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”
“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”
“Inference can’t happen without training.”
Precision is how many bits a number uses. FP16 (half precision) uses 16 bits; FP32 (single precision) uses 32. Mixed precision does most operations in half precision while keeping critical information in single precision. NVIDIA lists three benefits: less memory, less memory bandwidth, and much faster math, especially on GPUs with Tensor Cores.
What NVIDIA says (3)
“First, they require less memory, enabling the training and deployment of larger neural networks.”
“Second, they require less memory bandwidth which speeds up data transfer operations.”
“Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”
NVIDIA explains that training repeats a loop: the error is propagated back through the layers, the model adjusts its weights and tries again. Over many iterations it converges on the correct weights. NVIDIA calls training 'a monster when it comes to consuming compute.'
What NVIDIA says (2)
“Over many iterations, it converges on the correct weights, enabling it to reliably suggest accurate, useful code completions or translations.”
“The challenge is, it’s also a monster when it comes to consuming compute.”
Latency is the time from a request to its response. Throughput is how much work is completed per unit of time. NVIDIA says inference applies learned knowledge in real time. Low-latency, high-performance inference makes real-time interaction possible. NVIDIA adds that inference optimization lets enterprises tune throughput and latency to meet their service-level agreements (SLAs).
What NVIDIA says (2)
“This is only possible with low-latency, high-performance inference.”
“That’s why inference optimization techniques are so critical: they allow enterprises to optimize throughput and latency so they can meet their service level agreements across a variety of use cases.”
A large language model (LLM) generates text one token at a time. To avoid recomputing earlier tokens, it stores their attention keys and values in a KV (key-value) cache. NVIDIA says the two main contributors to GPU memory for LLM inference are the model weights and the KV cache. With batching, each request's KV cache is allocated separately, which can use a lot of memory.
What NVIDIA says (2)
“In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”
“With batching, the KV cache of each of the requests in the batch must still be allocated separately, and can have a large memory footprint.”
Key terms
- Training: The process of feeding data through a model and adjusting its weights until its outputs are accurate.
- Inference: Using a trained model to make predictions or generate outputs on new data.
- Mixed precision: Training with lower-precision number formats where safe, which saves memory and runs math faster on Tensor Cores.
- Latency: How long one request or one transfer takes from start to finish.
- Throughput: How much total work finishes per second, for example requests or samples per second.
- Batch: The group of samples a model processes in one step; batch size is how many samples are in that group.
- Gradient: The correction signal computed in training that says how much, and in which direction, to adjust each weight.
Try it
Sample question
Which statement best contrasts training and inference?
Show the answer
Answer: Training adjusts a model's weights by passing data through it and propagating errors back. Inference uses the trained model to make predictions on new data, often in real time.
Training is how a model learns. Data passes through the network, the error is propagated back through its layers, and the model adjusts its weights and tries again. A weight is one of the numbers inside the model that training tunes. Inference is when a trained model makes predictions on new data. NVIDIA notes inference cannot happen without training.
What NVIDIA says (3)
“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”
“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”
“Inference can’t happen without training.”
Practice 1.2 (5 questions) Full Essential AI Knowledge guide
← 1.1 The NVIDIA software stack · 1.3 AI, machine learning and deep learning →