5.2 Attention maps
What attention computes, how cross-attention keys shape attention maps, and attention as explanation.
Key points
Attention lets a model weigh other parts of the input when processing one part. An attention map is a picture of those weights, for example over image regions or words.
What NVIDIA says (1)
“Attention units follow these tags, calculating a kind of algebraic map of how each element relates to the others.”
Cross-attention lets image features attend to prompt words. Each word gets an attention map over the image. NVIDIA found the keys decide where those maps fall. VAE means variational autoencoder.
What NVIDIA says (1)
“Our main insight is that the key pathway of the cross-attention module in the diffusion model (the K matrix) controls the layout of the attention maps.”
Reading attention maps helps diagnose models. Leaking attention means the new word affects parts of the image it should not.
What NVIDIA says (1)
“existing techniques tend to overfit that component, causing the attention on”
Attention maps highlight which input regions or tokens mattered most. That makes a single decision easier to inspect.
What NVIDIA says (1)
“forcing the model itself to show its work.”
Key terms
- Cross-attention: Attention in which one sequence, such as text or image features, looks up information in another sequence.
- Attention map: A picture of attention weights showing which parts of the input a model focused on.
Sample question
What do attention units in a transformer compute?
Show the answer
Answer: A kind of map of how each element relates to the others
Attention lets a model weigh other parts of the input when processing one part. An attention map is a picture of those weights, for example over image regions or words.
What NVIDIA says (1)
“Attention units follow these tags, calculating a kind of algebraic map of how each element relates to the others.”
Practice 5.2 (4 questions) Full Data Analysis and Visualization guide
← 5.1 Insights from large datasets · 5.3 Graphs and charts →