4.1 Exploratory data analysis and descriptive statistics
describe(), null handling in reductions, and judging data gaps.
Key points
Descriptive statistics summarize the center, spread and shape of data. describe() is a fast first look during EDA. EDA means exploratory data analysis.
What NVIDIA says (2)
“For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”
“The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”
EDA comes before modeling. It checks what the data contains and whether it can be trusted.
What NVIDIA says (1)
“Now, you understand the following data characteristics: Data types Dimensions of the dataset Number of sources garnering the dataset Dataset update frequency However, you must still explore whether this data has major gaps, either with missing or invalid data inputs.”
Some missing data is acceptable. The rate of missing data tells you whether the source still represents reality. EDA means exploratory data analysis.
What NVIDIA says (2)
“To evaluate the rate of missing data, compare the number of readings to the expected number of readings.”
“Some missing data is acceptable, but if the stations went down too often, the data could be misrepresentative of true conditions throughout the year.”
A reduction turns a column into one number. Knowing how nulls are handled avoids silently biased statistics. NA means not available (missing).
What NVIDIA says (1)
“By default it’s value is set to True , we can change it to False to preserve NA values.”
Key terms
- Null (missing value): An empty entry where a value is unknown or absent.
- Exploratory data analysis: A first open-ended look at a dataset to learn its shape, gaps and patterns.
- Descriptive statistics: Summary numbers such as count, mean, standard deviation and percentiles.
Sample question
Which statistics does cuDF's describe() report for a numeric column?
Show the answer
Answer: count, mean, std, min, max and the 25th, 50th and 75th percentiles
Descriptive statistics summarize the center, spread and shape of data. describe() is a fast first look during EDA. EDA means exploratory data analysis.
What NVIDIA says (2)
“For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”
“The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”
Practice 4.1 (4 questions) Full Descriptive Analysis and Visualization guide
← 3.6 Reproducible pipelines with RAPIDS and Dask · 4.2 Visualization →