1.1 Joining and manipulating data with cuDF and pandas
What cuDF is, cudf.pandas zero-code-change acceleration, joins, group-bys, and why row loops are slow on GPUs.
Key points
A DataFrame is a table with named columns. cuDF holds DataFrames in GPU memory and offers the same style of API as pandas, the most common Python table library. API means application programming interface. CUDA means NVIDIA's parallel computing platform.
What NVIDIA says (1)
“cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”
A join (merge) combines two tables on matching key values. pandas keeps or sorts the key order; cuDF does not, because ordering costs time on a parallel GPU. Sort by the key or index after the merge when order matters.
What NVIDIA says (2)
“By contrast, cuDF’s default behavior is to return rows in a non-deterministic order to maximize performance.”
“Note that the dataframe order is not maintained , but may be restored post-merge by sorting by the index.”
cudf.pandas is cuDF's pandas accelerator mode. You keep your pandas code; supported operations run on the GPU, and anything else falls back to normal pandas. API means application programming interface. SQL means Structured Query Language.
What NVIDIA says (2)
“It supports 100% of the Pandas API , using the GPU for supported operations, and automatically falling back to pandas for other operations.”
“Nothing changes, not even your import statements, when going from CPU to GPU.”
cudf.pandas wraps each object in a proxy, a stand-in that can point to a GPU or a CPU copy. It tries the GPU first and falls back to the CPU only for the failing step.
What NVIDIA says (2)
“Attribute lookups and method calls are first attempted on the GPU (copying from CPU if necessary).”
“If that fails, the operation is attempted on the CPU (copying from GPU if necessary).”
GPUs are built to process many values at once, not one at a time. A vectorized function works on a whole column in one call.
What NVIDIA says (2)
“This is because iterating over data that resides on the GPU will yield extremely poor performance, as GPUs are optimized for highly parallel operations rather than sequential operations.”
“If you absolutely must iterate, copy the data from GPU to CPU by using .to_arrow() or .to_pandas() , then convert the result back to GPU using a Series , DataFrame or Index constructor.”
cuDF mirrors common pandas calls such as read_csv, groupby and agg. For code like this, swapping the import is enough. API means application programming interface. CUDA means NVIDIA's parallel computing platform.
What NVIDIA says (1)
“With RAPIDS, all you have to do to use this code to run on a GPU and enjoy the interactive querying of data is to change the import statement.”
Key terms
- pandas: A Python library for working with tables of data, called DataFrames.
- cuDF: A GPU DataFrame library with a pandas-style API for loading, joining, aggregating and filtering data.
- cudf.pandas: A cuDF mode that runs existing pandas code on the GPU without code changes and falls back to the CPU when needed.
- CPU fallback: Running an operation on the CPU when the GPU path does not support it, at the cost of extra data copies.
- Join (merge): Combining two tables by matching values in a key column.
- Group-by aggregation: Splitting rows into groups by a key and computing a summary, such as a sum, for each group.
Sample question
What is cuDF?
Show the answer
Answer: A Python GPU DataFrame library for loading, joining, aggregating and filtering tabular data with a pandas-style API
A DataFrame is a table with named columns. cuDF holds DataFrames in GPU memory and offers the same style of API as pandas, the most common Python table library. API means application programming interface. CUDA means NVIDIA's parallel computing platform.
What NVIDIA says (1)
“cuDF is a Python GPU DataFrame library (built on the Apache Arrow columnar memory format) for loading, joining, aggregating, filtering, and otherwise manipulating tabular data using a DataFrame style API in the style of pandas”
Practice 1.1 (6 questions) Full Data Manipulation and Preparation guide