1.7 Reading research and spotting trends
The ideas behind modern LLMs and the trends NVIDIA highlights.
Key points
Attention (self-attention) lets a model learn how even distant elements in a sequence relate to each other. NVIDIA's explainer credits the 2017 paper with defining transformers, and NVIDIA's exam page lists “Attention Is All You Need” as suggested reading.
What NVIDIA says (3)
“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.”
“Attention Is All You Need”
“one of eight co-authors of the 2017 paper that defined transformers.”
NVIDIA describes foundation models as neural networks trained on massive unlabeled datasets that, with a little fine-tuning, handle jobs from translating text to analyzing medical images.
What NVIDIA says (2)
“Foundation models are AI neural networks trained on massive unlabeled datasets to handle a wide variety of jobs from translating text to analyzing medical images.”
“With a little fine-tuning, foundation models can handle jobs from translating text to analyzing medical images to performing agent-based behaviors.”
NVIDIA's MoE glossary explains that a learned routing mechanism sparsely selects which subnetworks participate instead of running every parameter on every step.
What NVIDIA says (1)
“Instead of running every parameter on every step, a learned routing mechanism sparsely selects which subnetworks should participate, allowing the model to grow capacity without paying the full compute cost.”
Key terms
- Transformer: A neural-network architecture that uses attention to learn how parts of a sequence relate to each other.
- Attention: A mechanism that lets a model weigh how much each element of the input depends on every other element.
- Foundation model: A large model trained on massive unlabeled data that can be adapted to many tasks.
- Mixture of experts: A model made of many specialized sub-models (experts), where a router sends each token to only a few of them.
Sample question
The 2017 paper that introduced the Transformer is listed as suggested reading for NCA-GENL. What mechanism did it make central?
Show the answer
Answer: Attention (self-attention), which learns how elements in a sequence relate to each other
Attention (self-attention) lets a model learn how even distant elements in a sequence relate to each other. NVIDIA's explainer credits the 2017 paper with defining transformers, and NVIDIA's exam page lists “Attention Is All You Need” as suggested reading.
What NVIDIA says (3)
“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.”
“Attention Is All You Need”
“one of eight co-authors of the 2017 paper that defined transformers.”
Practice 1.7 (3 questions) Full Core Machine Learning and AI Knowledge guide