1.2 Training vs inference

NCA-AIIO · Essential AI Knowledge (38% of the exam) · Official objective: “Compare and contrast training and inference architecture requirements and considerations.”

What training and inference each do, and why they need different resources.

Key points

  1. Training is how a model learns. Data passes through the network, the error is propagated back through its layers, and the model adjusts its weights and tries again. A weight is one of the numbers inside the model that training tunes. Inference is when a trained model makes predictions on new data. NVIDIA notes inference cannot happen without training.

    What NVIDIA says (3)

    “Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”

    — What's the Difference Between Deep Learning Training and Inference?

    “Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”

    — What's the Difference Between Deep Learning Training and Inference?

    “Inference can’t happen without training.”

    — What's the Difference Between Deep Learning Training and Inference?

  2. Precision is how many bits a number uses. FP16 (half precision) uses 16 bits; FP32 (single precision) uses 32. Mixed precision does most operations in half precision while keeping critical information in single precision. NVIDIA lists three benefits: less memory, less memory bandwidth, and much faster math, especially on GPUs with Tensor Cores.

    What NVIDIA says (3)

    “First, they require less memory, enabling the training and deployment of larger neural networks.”

    — NVIDIA Mixed Precision Training Guide

    “Second, they require less memory bandwidth which speeds up data transfer operations.”

    — NVIDIA Mixed Precision Training Guide

    “Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”

    — NVIDIA Mixed Precision Training Guide

  3. NVIDIA explains that training repeats a loop: the error is propagated back through the layers, the model adjusts its weights and tries again. Over many iterations it converges on the correct weights. NVIDIA calls training 'a monster when it comes to consuming compute.'

    What NVIDIA says (2)

    “Over many iterations, it converges on the correct weights, enabling it to reliably suggest accurate, useful code completions or translations.”

    — What's the Difference Between Deep Learning Training and Inference?

    “The challenge is, it’s also a monster when it comes to consuming compute.”

    — What's the Difference Between Deep Learning Training and Inference?

  4. Latency is the time from a request to its response. Throughput is how much work is completed per unit of time. NVIDIA says inference applies learned knowledge in real time. Low-latency, high-performance inference makes real-time interaction possible. NVIDIA adds that inference optimization lets enterprises tune throughput and latency to meet their service-level agreements (SLAs).

    What NVIDIA says (2)

    “This is only possible with low-latency, high-performance inference.”

    — What is AI Inference?

    “That’s why inference optimization techniques are so critical: they allow enterprises to optimize throughput and latency so they can meet their service level agreements across a variety of use cases.”

    — What's the Difference Between Deep Learning Training and Inference?

  5. A large language model (LLM) generates text one token at a time. To avoid recomputing earlier tokens, it stores their attention keys and values in a KV (key-value) cache. NVIDIA says the two main contributors to GPU memory for LLM inference are the model weights and the KV cache. With batching, each request's KV cache is allocated separately, which can use a lot of memory.

    What NVIDIA says (2)

    “In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”

    — Mastering LLM Techniques: Inference Optimization

    “With batching, the KV cache of each of the requests in the batch must still be allocated separately, and can have a large memory footprint.”

    — Mastering LLM Techniques: Inference Optimization

Key terms

Try it

Sample question

Which statement best contrasts training and inference?

Show the answer

Answer: Training adjusts a model's weights by passing data through it and propagating errors back. Inference uses the trained model to make predictions on new data, often in real time.

Training is how a model learns. Data passes through the network, the error is propagated back through its layers, and the model adjusts its weights and tries again. A weight is one of the numbers inside the model that training tunes. Inference is when a trained model makes predictions on new data. NVIDIA notes inference cannot happen without training.

What NVIDIA says (3)

“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”

— What's the Difference Between Deep Learning Training and Inference?

“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”

— What's the Difference Between Deep Learning Training and Inference?

“Inference can’t happen without training.”

— What's the Difference Between Deep Learning Training and Inference?

Practice 1.2 (5 questions) Full Essential AI Knowledge guide

← 1.1 The NVIDIA software stack · 1.3 AI, machine learning and deep learning →