1.4 Why AI took off
The factors behind the recent jump in AI: GPUs, transformers and pretrained models.
Key points
Parallel computing means doing many calculations at the same time. NVIDIA says the parallelism of deep learning maps naturally to GPUs. That gives a significant speedup over CPU-only training and made GPUs the platform of choice for large neural networks. NVIDIA notes GPU-accelerated frameworks can cut training from days and weeks to hours and days.
What NVIDIA says (2)
“This parallelism maps naturally to GPUs , providing a significant computation speedup over CPU-only training and making them the platform of choice for training large, complex neural network-based systems.”
“researchers and data scientists can significantly speed up deep learning training that could otherwise take days and weeks to just hours and days.”
Labeled data has the right answer attached to each example, which takes people time and money to create. Self-supervised learning finds its training signal in the data itself. NVIDIA says that before transformers, users had to train with large, labeled datasets that were costly and time-consuming to produce. NVIDIA's CEO is quoted: transformers made self-supervised learning possible, and AI jumped to warp speed.
What NVIDIA says (2)
“Before transformers arrived, users had to train neural networks with large, labeled datasets that were costly and time-consuming to produce.”
“Transformers made self-supervised learning possible, and AI jumped to warp speed”
A pretrained model is a deep learning model trained on large datasets to do a specific task. Fine-tuning is extra training that adapts it to a narrower need. NVIDIA says a pretrained model can be used as is or customized to suit application requirements across multiple industries.
What NVIDIA says (3)
“A pretrained AI model is a deep learning model that’s trained on large datasets to accomplish a specific task, and it can be used as is or customized to suit application requirements across multiple industries.”
“You can use the pretrained models for inference or fine-tune them with transfer learning.”
“It can be used as is or further fine-tuned to fit an application’s specific needs.”
Key terms
- Deep learning: A subset of machine learning that uses neural networks with many layers to learn patterns directly from data such as images or text.
- Transformer: A neural network design that uses attention to relate every part of an input to every other part, and runs well in parallel.
- Pretrained model: A model already trained on a large dataset that you can use as is or fine-tune for your task.
- Accelerated computing: Using specialized hardware such as GPUs to speed up demanding work through parallel processing.
- Central processing unit: The general-purpose processor in a server: a few fast cores with large caches, built to run a few threads quickly.
Try it
Sample question
Which factor does NVIDIA credit with making GPUs the platform of choice for training large neural networks?
Show the answer
Answer: Neural network math is highly parallel, which maps naturally to GPUs and gives a large speedup over CPU-only training.
Parallel computing means doing many calculations at the same time. NVIDIA says the parallelism of deep learning maps naturally to GPUs. That gives a significant speedup over CPU-only training and made GPUs the platform of choice for large neural networks. NVIDIA notes GPU-accelerated frameworks can cut training from days and weeks to hours and days.
What NVIDIA says (2)
“This parallelism maps naturally to GPUs , providing a significant computation speedup over CPU-only training and making them the platform of choice for training large, complex neural network-based systems.”
“researchers and data scientists can significantly speed up deep learning training that could otherwise take days and weeks to just hours and days.”
Practice 1.4 (3 questions) Full Essential AI Knowledge guide
← 1.3 AI, machine learning and deep learning · 1.5 AI use cases and industries →