5.3 CPU versus GPU workloads and memory transfers

NCA-ADS · Foundations of Accelerated Data Science (12% of the exam) · Official objective: “CPU vs. GPU workloads and memory transfer optimization”

When small data is faster on the CPU, minimizing host-device copies and pinned memory.

Key points

  1. Moving data between CPU and GPU memory takes time. With little data, that fixed cost dominates. GPUs shine on many rows.

    What NVIDIA says (2)

    “Small data sizes may be slower on GPU than CPU, because of the cost of data transfers.”

    — cuDF: cudf.pandas FAQ and Known Issues

    “cuDF achieves the highest performance with many rows of data.”

    — cuDF: cudf.pandas FAQ and Known Issues

  2. The host is the CPU side; the device is the GPU side. The link between them is much slower than GPU memory, so keep data on the GPU.

    What NVIDIA says (2)

    “Minimize the amount of data transferred between host and device when possible, even if that means running kernels on the GPU that get little or no speed-up compared to running them on the host CPU.”

    — How to Optimize Data Transfers in CUDA C/C++

    “Batching many small transfers into one larger transfer performs much better because it eliminates most of the per-transfer overhead.”

    — How to Optimize Data Transfers in CUDA C/C++

  3. Pageable memory can be moved by the operating system; pinned memory cannot. Copying from pinned memory skips an extra staging copy.

    What NVIDIA says (2)

    “Higher bandwidth is possible between the host and the device when using page-locked (or “pinned”) memory.”

    — How to Optimize Data Transfers in CUDA C/C++

    “The GPU cannot access data directly from pageable host memory, so when a data transfer from pageable host memory to device memory is invoked, the CUDA driver must first allocate a temporary page-locked, or “pinned”, host array, copy the host data to the pinned array, and then transfer the data f”

    — How to Optimize Data Transfers in CUDA C/C++

Key terms

Try it

Sample question

A table has 2,000 rows. cudf.pandas runs it slower than plain pandas. Why?

Show the answer

Answer: Small data can be slower on the GPU because data transfer costs outweigh the parallel speedup

Moving data between CPU and GPU memory takes time. With little data, that fixed cost dominates. GPUs shine on many rows.

What NVIDIA says (2)

“Small data sizes may be slower on GPU than CPU, because of the cost of data transfers.”

— cuDF: cudf.pandas FAQ and Known Issues

“cuDF achieves the highest performance with many rows of data.”

— cuDF: cudf.pandas FAQ and Known Issues

Practice 5.3 (3 questions) Full Foundations of Accelerated Data Science guide

← 5.2 Why GPUs speed up data science · 5.4 The end-to-end data science workflow →