4.1 Exploratory data analysis and descriptive statistics

NCA-ADS · Descriptive Analysis and Visualization (13% of the exam) · Official objective: “Exploratory data analysis (EDA) and descriptive statistics”

describe(), null handling in reductions, and judging data gaps.

Key points

  1. Descriptive statistics summarize the center, spread and shape of data. describe() is a fast first look during EDA. EDA means exploratory data analysis.

    What NVIDIA says (2)

    “For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”

    — cuDF API: DataFrame.describe

    “The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”

    — cuDF API: DataFrame.describe

  2. EDA comes before modeling. It checks what the data contains and whether it can be trusted.

    What NVIDIA says (1)

    “Now, you understand the following data characteristics: Data types Dimensions of the dataset Number of sources garnering the dataset Dataset update frequency However, you must still explore whether this data has major gaps, either with missing or invalid data inputs.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  3. Some missing data is acceptable. The rate of missing data tells you whether the source still represents reality. EDA means exploratory data analysis.

    What NVIDIA says (2)

    “To evaluate the rate of missing data, compare the number of readings to the expected number of readings.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

    “Some missing data is acceptable, but if the stations went down too often, the data could be misrepresentative of true conditions throughout the year.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  4. A reduction turns a column into one number. Knowing how nulls are handled avoids silently biased statistics. NA means not available (missing).

    What NVIDIA says (1)

    “By default it’s value is set to True , we can change it to False to preserve NA values.”

    — cuDF: Working with missing data

Key terms

Sample question

Which statistics does cuDF's describe() report for a numeric column?

Show the answer

Answer: count, mean, std, min, max and the 25th, 50th and 75th percentiles

Descriptive statistics summarize the center, spread and shape of data. describe() is a fast first look during EDA. EDA means exploratory data analysis.

What NVIDIA says (2)

“For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”

— cuDF API: DataFrame.describe

“The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”

— cuDF API: DataFrame.describe

Practice 4.1 (4 questions) Full Descriptive Analysis and Visualization guide

← 3.6 Reproducible pipelines with RAPIDS and Dask · 4.2 Visualization →