1.7 Efficient storage with Parquet and modern frameworks

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “Efficient processing and storage with Parquet and modern frameworks”

Column pruning, filter pushdown, partitioned writes, GPUDirect Storage and the Polars GPU engine.

Key points

  1. Parquet is a columnar file format: each column is stored separately. So a reader can load only the columns you ask for. CSV means comma-separated values. I/O means input/output.

    What NVIDIA says (1)

    “columns list, default None If not None, only these columns will be read.”

    — cuDF API: read_parquet

  2. A row group is a block of rows inside a Parquet file. Each row group stores min/max statistics, so the reader can skip blocks that cannot match.

    What NVIDIA says (1)

    “If not None, specifies a filter predicate used to filter out row groups using statistics stored for each row group as Parquet metadata.”

    — cuDF API: read_parquet

  3. Partitioning a dataset by a column puts each value's rows in its own folder. Readers can then skip folders they do not need.

    What NVIDIA says (1)

    “partition_cols list, optional, default None Column names by which to partition the dataset Columns are partitioned in the order they are given”

    — cuDF API: DataFrame.to_parquet

  4. GPUDirect Storage lets data flow between storage and GPU memory without an extra copy through CPU memory. cuDF uses it for readers such as read_parquet when available. CSV means comma-separated values.

    What NVIDIA says (1)

    “cuDF leverages the KvikIO library for high-performance I/O features, such as parallel I/O operations and NVIDIA Magnum IO GPUDirect Storage (GDS).”

    — cuDF: Input / Output

  5. Polars is a DataFrame library with a lazy API, which builds a query plan before running it. cuDF provides a GPU engine for that plan. API means application programming interface.

    What NVIDIA says (2)

    “cuDF provides GPU-accelerated execution engines for Python users of the Polars Lazy API.”

    — cuDF: Polars GPU engine

    “If it is not, the execution transparently falls back to the standard Polars engine and runs on the CPU.”

    — cuDF: Polars GPU engine

Key terms

Sample question

A Parquet file has 200 columns but your analysis needs 5. How do you cut I/O with cudf.read_parquet?

Show the answer

Answer: Pass columns=[...] so only those columns are read

Parquet is a columnar file format: each column is stored separately. So a reader can load only the columns you ask for. CSV means comma-separated values. I/O means input/output.

What NVIDIA says (1)

“columns list, default None If not None, only these columns will be read.”

— cuDF API: read_parquet

Practice 1.7 (5 questions) Full Data Manipulation and Preparation guide

← 1.6 Dimensionality reduction and sampling · 2.1 Training on GPUs with cuML and XGBoost →