Descriptive Analysis and Visualization
13% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management
4.1 Exploratory data analysis and descriptive statistics
describe(), null handling in reductions, and judging data gaps.
Key points
Descriptive statistics summarize the center, spread and shape of data. describe() is a fast first look during EDA. EDA means exploratory data analysis.
What NVIDIA says (2)
“For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”
“The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”
EDA comes before modeling. It checks what the data contains and whether it can be trusted.
What NVIDIA says (1)
“Now, you understand the following data characteristics: Data types Dimensions of the dataset Number of sources garnering the dataset Dataset update frequency However, you must still explore whether this data has major gaps, either with missing or invalid data inputs.”
Some missing data is acceptable. The rate of missing data tells you whether the source still represents reality. EDA means exploratory data analysis.
What NVIDIA says (2)
“To evaluate the rate of missing data, compare the number of readings to the expected number of readings.”
“Some missing data is acceptable, but if the stations went down too often, the data could be misrepresentative of true conditions throughout the year.”
A reduction turns a column into one number. Knowing how nulls are handled avoids silently biased statistics. NA means not available (missing).
What NVIDIA says (1)
“By default it’s value is set to True , we can change it to False to preserve NA values.”
Key terms: Null (missing value) Exploratory data analysis Descriptive statistics
4.2 Visualization
Why visualize, Datashader for millions of points and interactive dashboards.
Key points
Visualization means showing data as charts. The eye can spot unusual points and shapes quickly.
What NVIDIA says (1)
“Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns not easily surfaced by purely analytical methods.”
Overplotting happens when too many points overlap and hide the pattern. Datashader aggregates points into pixels to show density.
What NVIDIA says (2)
“The Datashader library directly supports cuDF and can rapidly render over millions of aggregated points.”
“Datapoint rendering displaying high-resolution patterns is precisely what Datashader is designed for.”
Dash builds interactive web dashboards in Python. With cuDF behind it, the app can stay fast on large data.
What NVIDIA says (1)
“Plotly Dash enables data scientists to recast complex data and machine learning workflows as more accessible web applications.”
Precomputed aggregations are summaries prepared in advance. GPU speed lets the dashboard compute them on the fly as users interact.
What NVIDIA says (1)
“The use of Plotly’s Dash, RAPIDS, and Data shader allows users to build viz dashboards that both render datasets of 300 million+ rows and remain highly interactive without the need for precomputed aggregations.”
Key terms: Datashader Plotly Dash
4.3 Choosing the right plot
Histograms, heat maps, map plots and cross-filtering.
Key points
A histogram counts how many values fall into each range. It shows the shape of a distribution, such as most trips being short.
What NVIDIA says (1)
“An hvPlot histogram of trip durations generated with the Divvy dataset In this instance, the vast majority of bike trips appear under 20 minutes.”
A heat map colors a grid by a value across two categories. It makes daily and weekly patterns easy to see.
What NVIDIA says (1)
“An hvPlot heat map showing trips by hour and day of week, per month Adding a widget for interactivity enables scrubbing through the months to search for patterns over a full year (Figure 2).”
A hexbin chart groups nearby points into hexagon cells and colors them by count. It keeps maps readable for many points.
What NVIDIA says (1)
“Figure 3 shows the hexbin chart that aggregates trip start and ending locations to a manageable amount, verifying that the data is accurate to the bike share system map.”
Cross-filtering links charts so a selection in one filters the others. It replaces writing many separate groupby and query calls.
What NVIDIA says (1)
“Instead of creating several individual group by and query operations, a cuxfilter dashboard can simply cross-link numerous charts to quickly find patterns or anomalies (Figure 5).”
Key terms: cuxfilter Histogram Heat map Cross-filtering
4.4 Hypothesis testing and significance
p-values, t-tests, confidence intervals and sample size.
Key points
A hypothesis test asks whether an observed difference could easily happen by chance. The p-value is that chance under the 'no difference' assumption; a small p-value means significance.
What NVIDIA says (1)
“The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”
A confidence interval is a range that likely contains the true value. More trials shrink it; too few trials leave comparisons inconclusive.
What NVIDIA says (2)
“A single success rate on N rollouts tells you almost nothing about how confident you should be in a policy’s true performance.”
“Most published benchmarks do not run a sufficient number of rollouts to achieve statistical significance when comparing the performance of two policies.”
The width of a confidence interval shrinks roughly with the square root of the sample size. So precision gets expensive.
What NVIDIA says (1)
“Narrowing the confidence interval from 10 to 2 percentage points requires roughly 15x more rollouts (70 to 1,030).”
Key terms: p-value Confidence interval
4.5 Patterns, trends and relationships
Pearson and Spearman correlation, rolling means and changing trends.
Key points
Correlation measures how two variables move together. Causation needs more evidence, such as a controlled experiment.
What NVIDIA says (2)
“pearson : Standard correlation coefficient spearman : Spearman rank correlation”
“However, it’s important to remember that correlation and causation are two different things.”
Monotonic means one variable tends to rise when the other rises, not always at the same rate. Spearman uses ranks, so it captures that without assuming a line.
What NVIDIA says (1)
“Method used to compute correlation: pearson : Standard correlation coefficient spearman : Spearman rank correlation”
A trend is the long-run direction of a series. A rolling window averages nearby values to smooth out noise.
What NVIDIA says (1)
“Parameters : window int, offset or a BaseIndexer subclass Size of the window, i.e., the number of observations used to calculate the statistic.”
A linear model assumes a straight-line relationship. When the pattern curves, a nonlinear model can describe it better.
What NVIDIA says (2)
“While there is a strong relationship between population and time, the relationship is not linear because various factors influence changes from year to year.”
“Nonlinear regression can estimate models with arbitrary relationships between independent and dependent variables.”
Key terms: Correlation Rolling mean