1.8 GPU vs CPU architecture

NCA-AIIO · Essential AI Knowledge (38% of the exam) · Official objective: “Compare and contrast GPU and CPU architectures.”

Why GPUs, with many cores running in parallel, suit AI better than CPUs for the heavy math.

Key points

  1. NVIDIA's CUDA guide says a CPU is designed to run a serial sequence of operations as fast as possible. A GPU is designed to run thousands of threads in parallel, trading lower single-thread performance for much greater total throughput. GPUs devote more transistors to data processing, while CPUs devote more to caching and flow control.

    What NVIDIA says (2)

    “a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”

    — CUDA Programming Guide: Introduction

    “GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”

    — CUDA Programming Guide: Introduction

  2. A core is a processing unit. A thread is one sequence of instructions. NVIDIA says a CPU has a few cores with lots of cache memory, handling a few software threads at a time. A GPU has hundreds of cores that handle thousands of threads simultaneously.

    What NVIDIA says (5)

    “GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”

    — CUDA Programming Guide: Introduction

    “a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”

    — CUDA Programming Guide: Introduction

    “Architecturally, the CPU is composed of just a few cores with lots of cache memory that can handle a few software threads at a time.”

    — What's the Difference Between a CPU and a GPU?

    “a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”

    — CUDA Programming Guide: Introduction

    “In contrast, a GPU is composed of hundreds of cores that can handle thousands of threads simultaneously.”

    — What's the Difference Between a CPU and a GPU?

  3. A tensor is a multi-dimensional array of numbers, the basic data type of deep learning. NVIDIA says deep learning is accelerated by dedicated Tensor Cores in its GPUs. Its mixed-precision guide notes math runs much faster in reduced precision on GPUs with Tensor Core support.

    What NVIDIA says (2)

    “Tensor Cores were introduced in the NVIDIA Volta™ GPU architecture to accelerate matrix multiply and accumulate operations for machine learning and scientific applications.”

    — GPU Performance Background User's Guide

    “Third, math operations run much faster in reduced precision, especially on GPUs with Tensor Core support for that precision.”

    — NVIDIA Mixed Precision Training Guide

  4. NVIDIA defines accelerated computing as using specialized hardware to dramatically speed up work with parallel processing. It offloads demanding work that can bog down CPUs, which typically run tasks one after another. The CPU still runs the parts of the program that are serial.

    What NVIDIA says (4)

    “A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”

    — CUDA Programming Guide: Introduction

    “Accelerated computing is the use of specialized hardware to dramatically speed up work, using parallel processing that bundles frequently occurring tasks.”

    — What Is Accelerated Computing?

    “a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”

    — CUDA Programming Guide: Introduction

    “It offloads demanding work that can bog down CPUs, processors that typically execute tasks in serial fashion.”

    — What Is Accelerated Computing?

  5. Memory bandwidth is how fast data moves between memory and the processor. NVIDIA's CUDA guide says a GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.

    What NVIDIA says (1)

    “A GPU provides much higher instruction throughput and memory bandwidth than a CPU within a similar price and power envelope.”

    — CUDA Programming Guide: Introduction

Key terms

Try it

Sample question

What trade-off does a GPU make compared with a CPU?

Show the answer

Answer: It gives up some single-thread performance to run thousands of threads in parallel for much greater total throughput.

NVIDIA's CUDA guide says a CPU is designed to run a serial sequence of operations as fast as possible. A GPU is designed to run thousands of threads in parallel, trading lower single-thread performance for much greater total throughput. GPUs devote more transistors to data processing, while CPUs devote more to caching and flow control.

What NVIDIA says (2)

“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”

— CUDA Programming Guide: Introduction

“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”

— CUDA Programming Guide: Introduction

Practice 1.8 (5 questions) Full Essential AI Knowledge guide

← 1.7 The AI development and deployment life cycle · 2.1 Hardware for training workloads →