5.1 Insights from large datasets

NCA-GENM · Data Analysis and Visualization (10% of the exam) · Official objective: “Awareness of the process of extracting insights from large datasets using data mining, data visualization, and similar techniques.”

Scaling exploration with cuDF, finding gaps, clustering and early visualization.

Key points

  1. pandas runs on one CPU core and slows past 1-2 GB. cuDF parallelizes the same kind of operations across GPU cores.

    What NVIDIA says (2)

    “pandas is designed to run on a single core and starts slowing down when data size hits 1-2 GB”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

    “For that middle ground of 2-10 GB, RAPIDS cuDF is the Goldilocks solution that is just right.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  2. Exploratory data analysis (EDA) is the first look at data to learn its shape and problems. Gaps affect whether the data can be relied on alone.

    What NVIDIA says (1)

    “However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  3. Clustering groups similar items without labels. It can reveal unknown patterns, such as customer segments or image themes.

    What NVIDIA says (1)

    “Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”

    — What is Machine Learning and Why Does It Matter?

  4. Anscombe's quartet shows datasets with the same statistics but very different shapes. Looking early catches such surprises.

    What NVIDIA says (1)

    “Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms

Sample question

Your pandas exploration of a 5 GB multimodal metadata table is slow. What does NVIDIA suggest for that size?

Show the answer

Answer: RAPIDS cuDF, which runs a pandas-like API on the GPU

pandas runs on one CPU core and slows past 1-2 GB. cuDF parallelizes the same kind of operations across GPU cores.

What NVIDIA says (2)

“pandas is designed to run on a single core and starts slowing down when data size hits 1-2 GB”

— Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

“For that middle ground of 2-10 GB, RAPIDS cuDF is the Goldilocks solution that is just right.”

— Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

Practice 5.1 (4 questions) Full Data Analysis and Visualization guide

← 4.5 CLIP for text-to-image · 5.2 Attention maps →