1.1 Joining and manipulating data with cuDF and pandas

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “Data integration, joining, and manipulation using NVIDIA cuDF and pandas”

What cuDF is, cudf.pandas zero-code-change acceleration, joins, group-bys, and why row loops are slow on GPUs.

Key points

  1. A DataFrame is a table with named columns. cuDF holds DataFrames in GPU memory and offers the same style of API as pandas, the most common Python table library. API means application programming interface. CUDA means NVIDIA's parallel computing platform.

    What NVIDIA says (1)

    “cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”

    — cuDF: 10 Minutes to cuDF and Dask cuDF

  2. A join (merge) combines two tables on matching key values. pandas keeps or sorts the key order; cuDF does not, because ordering costs time on a parallel GPU. Sort by the key or index after the merge when order matters.

    What NVIDIA says (2)

    “By contrast, cuDF’s default behavior is to return rows in a non-deterministic order to maximize performance.”

    — cuDF: Comparison of cuDF and pandas

    “Note that the dataframe order is not maintained , but may be restored post-merge by sorting by the index.”

    — cuDF: 10 Minutes to cuDF and Dask cuDF

  3. cudf.pandas is cuDF's pandas accelerator mode. You keep your pandas code; supported operations run on the GPU, and anything else falls back to normal pandas. API means application programming interface. SQL means Structured Query Language.

    What NVIDIA says (2)

    “It supports 100% of the Pandas API , using the GPU for supported operations, and automatically falling back to pandas for other operations.”

    — cuDF: cudf.pandas

    “Nothing changes, not even your import statements, when going from CPU to GPU.”

    — cuDF: cudf.pandas

  4. cudf.pandas wraps each object in a proxy, a stand-in that can point to a GPU or a CPU copy. It tries the GPU first and falls back to the CPU only for the failing step.

    What NVIDIA says (2)

    “Attribute lookups and method calls are first attempted on the GPU (copying from CPU if necessary).”

    — cuDF: How cudf.pandas works

    “If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”

    — cuDF: How cudf.pandas works

  5. GPUs are built to process many values at once, not one at a time. A vectorized function works on a whole column in one call.

    What NVIDIA says (2)

    “This is because iterating over data that resides on the GPU will yield extremely poor performance, as GPUs are optimized for highly parallel operations rather than sequential operations.”

    — cuDF: Comparison of cuDF and pandas

    “If you absolutely must iterate, copy the data from GPU to CPU by using .to_arrow() or .to_pandas() , then convert the result back to GPU using a Series , DataFrame or Index constructor.”

    — cuDF: Comparison of cuDF and pandas

  6. cuDF mirrors common pandas calls such as read_csv, groupby and agg. For code like this, swapping the import is enough. API means application programming interface. CUDA means NVIDIA's parallel computing platform.

    What NVIDIA says (1)

    “With RAPIDS, all you have to do to use this code to run on a GPU and enjoy the interactive querying of data is to change the import statement.”

    — Beginner's Guide to GPU-Accelerated DataFrames for pandas Users

Key terms

Sample question

What is cuDF?

Show the answer

Answer: A Python GPU DataFrame library for loading, joining, aggregating and filtering tabular data with a pandas-style API

A DataFrame is a table with named columns. cuDF holds DataFrames in GPU memory and offers the same style of API as pandas, the most common Python table library. API means application programming interface. CUDA means NVIDIA's parallel computing platform.

What NVIDIA says (1)

“cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”

— cuDF: 10 Minutes to cuDF and Dask cuDF

Practice 1.1 (6 questions) Full Data Manipulation and Preparation guide

1.2 Cleaning data and handling quality and governance →