1.8 GPU vs CPU architecture
Why GPUs, with many cores running in parallel, suit AI better than CPUs for the heavy math.
Key points
NVIDIA's CUDA guide says a CPU is designed to run a serial sequence of operations as fast as possible. A GPU is designed to run thousands of threads in parallel, trading lower single-thread performance for much greater total throughput. GPUs devote more transistors to data processing, while CPUs devote more to caching and flow control.
What NVIDIA says (2)
“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”
“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”
A core is a processing unit. A thread is one sequence of instructions. NVIDIA says a CPU has a few cores with lots of cache memory, handling a few software threads at a time. A GPU has hundreds of cores that handle thousands of threads simultaneously.
What NVIDIA says (5)
“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”
“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”
“Architecturally, the CPU is composed of just a few cores with lots of cache memory that can handle a few software threads at a time.”
“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”
“In contrast, a GPU is composed of hundreds of cores that can handle thousands of threads simultaneously.”
A tensor is a multi-dimensional array of numbers, the basic data type of deep learning. NVIDIA says deep learning is accelerated by dedicated Tensor Cores in its GPUs. Its mixed-precision guide notes math runs much faster in reduced precision on GPUs with Tensor Core support.
What NVIDIA says (2)
“Tensor Cores were introduced in the NVIDIA Volta™ GPU architecture to accelerate matrix multiply and accumulate operations for machine learning and scientific applications.”
“Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”
NVIDIA defines accelerated computing as using specialized hardware to dramatically speed up work with parallel processing. It offloads demanding work that can bog down CPUs, which typically run tasks one after another. The CPU still runs the parts of the program that are serial.
What NVIDIA says (4)
“A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”
“Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”
“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”
“It offloads demanding work that can bog down CPUs, processors that typically execute tasks in serial fashion.”
Memory bandwidth is how fast data moves between memory and the processor. NVIDIA's CUDA guide says a GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.
What NVIDIA says (1)
“A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”
Key terms
- Mixed precision: Training with lower-precision number formats where safe, which saves memory and runs math faster on Tensor Cores.
- Tensor Core: A specialized unit inside NVIDIA GPUs that speeds up the matrix math used in deep learning.
- Graphics processing unit: A processor with hundreds or thousands of cores that runs many threads in parallel for high total throughput.
- Accelerated computing: Using specialized hardware such as GPUs to speed up demanding work through parallel processing.
Try it
Sample question
What trade-off does a GPU make compared with a CPU?
Show the answer
Answer: It gives up some single-thread performance to run thousands of threads in parallel for much greater total throughput.
NVIDIA's CUDA guide says a CPU is designed to run a serial sequence of operations as fast as possible. A GPU is designed to run thousands of threads in parallel, trading lower single-thread performance for much greater total throughput. GPUs devote more transistors to data processing, while CPUs devote more to caching and flow control.
What NVIDIA says (2)
“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”
“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”
Practice 1.8 (5 questions) Full Essential AI Knowledge guide
← 1.7 The AI development and deployment life cycle · 2.1 Hardware for training workloads →