5.1 Insights from large datasets
Scaling exploration with cuDF, finding gaps, clustering and early visualization.
Key points
pandas runs on one CPU core and slows past 1-2 GB. cuDF parallelizes the same kind of operations across GPU cores.
What NVIDIA says (2)
“pandas is designed to run on a single core and starts slowing down when data size hits 1-2 GB”
“For that middle ground of 2-10 GB, RAPIDS cuDF is the Goldilocks solution that is just right.”
Exploratory data analysis (EDA) is the first look at data to learn its shape and problems. Gaps affect whether the data can be relied on alone.
What NVIDIA says (1)
“However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”
Clustering groups similar items without labels. It can reveal unknown patterns, such as customer segments or image themes.
What NVIDIA says (1)
“Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”
Anscombe's quartet shows datasets with the same statistics but very different shapes. Looking early catches such surprises.
What NVIDIA says (1)
“Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns”
Key terms
- pandas: The most popular Python library for working with tables of data.
- RAPIDS cuDF: A GPU dataframe library with a pandas-like API.
- Exploratory data analysis: A first look at data to learn its shape, gaps and patterns.
- Clustering: Grouping similar items without labels; a common unsupervised task.
Sample question
Your pandas exploration of a 5 GB multimodal metadata table is slow. What does NVIDIA suggest for that size?
Show the answer
Answer: RAPIDS cuDF, which runs a pandas-like API on the GPU
pandas runs on one CPU core and slows past 1-2 GB. cuDF parallelizes the same kind of operations across GPU cores.
What NVIDIA says (2)
“pandas is designed to run on a single core and starts slowing down when data size hits 1-2 GB”
“For that middle ground of 2-10 GB, RAPIDS cuDF is the Goldilocks solution that is just right.”
Practice 5.1 (4 questions) Full Data Analysis and Visualization guide