1.7 Efficient storage with Parquet and modern frameworks
Column pruning, filter pushdown, partitioned writes, GPUDirect Storage and the Polars GPU engine.
Key points
Parquet is a columnar file format: each column is stored separately. So a reader can load only the columns you ask for. CSV means comma-separated values. I/O means input/output.
What NVIDIA says (1)
“columns list, default None If not None, only these columns will be read.”
A row group is a block of rows inside a Parquet file. Each row group stores min/max statistics, so the reader can skip blocks that cannot match.
What NVIDIA says (1)
“If not None, specifies a filter predicate used to filter out row groups using statistics stored for each row group as Parquet metadata.”
Partitioning a dataset by a column puts each value's rows in its own folder. Readers can then skip folders they do not need.
What NVIDIA says (1)
“partition_cols list, optional, default None Column names by which to partition the dataset Columns are partitioned in the order they are given”
GPUDirect Storage lets data flow between storage and GPU memory without an extra copy through CPU memory. cuDF uses it for readers such as read_parquet when available. CSV means comma-separated values.
What NVIDIA says (1)
“cuDF leverages the KvikIO library for high-performance I/O features, such as parallel I/O operations and NVIDIA Magnum IO GPUDirect Storage (GDS).”
Polars is a DataFrame library with a lazy API, which builds a query plan before running it. cuDF provides a GPU engine for that plan. API means application programming interface.
What NVIDIA says (2)
“cuDF provides GPU-accelerated execution engines for Python users of the Polars Lazy API.”
“If it is not, the execution transparently falls back to the standard Polars engine and runs on the CPU.”
Key terms
- Polars GPU engine: A cuDF-powered engine that runs Polars lazy queries on the GPU and falls back to the CPU when needed.
- KvikIO: A library cuDF uses for fast, parallel file input and output, including GPUDirect Storage.
- Parquet: A columnar file format that stores data by column, so tools can read only the columns they need.
- Partitioned dataset: A dataset split into folders by column values, such as one folder per year.
Sample question
A Parquet file has 200 columns but your analysis needs 5. How do you cut I/O with cudf.read_parquet?
Show the answer
Answer: Pass columns=[...] so only those columns are read
Parquet is a columnar file format: each column is stored separately. So a reader can load only the columns you ask for. CSV means comma-separated values. I/O means input/output.
What NVIDIA says (1)
“columns list, default None If not None, only these columns will be read.”
Practice 1.7 (5 questions) Full Data Manipulation and Preparation guide
← 1.6 Dimensionality reduction and sampling · 2.1 Training on GPUs with cuML and XGBoost →