Descriptive Analysis and Visualization

13% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

4.1 Exploratory data analysis and descriptive statistics

Official objective: “Exploratory data analysis (EDA) and descriptive statistics”

describe(), null handling in reductions, and judging data gaps.

Key points

  1. Descriptive statistics summarize the center, spread and shape of data. describe() is a fast first look during EDA. EDA means exploratory data analysis.

    What NVIDIA says (2)

    “For numeric data, the result’s index will include count , mean , std , min , max as well as lower, 50 and upper percentiles.”

    — cuDF API: DataFrame.describe

    “The default is [.25, .5, .75] , which returns the 25th, 50th, and 75th percentiles.”

    — cuDF API: DataFrame.describe

  2. EDA comes before modeling. It checks what the data contains and whether it can be trusted.

    What NVIDIA says (1)

    “Now, you understand the following data characteristics: Data types Dimensions of the dataset Number of sources garnering the dataset Dataset update frequency However, you must still explore whether this data has major gaps, either with missing or invalid data inputs.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  3. Some missing data is acceptable. The rate of missing data tells you whether the source still represents reality. EDA means exploratory data analysis.

    What NVIDIA says (2)

    “To evaluate the rate of missing data, compare the number of readings to the expected number of readings.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

    “Some missing data is acceptable, but if the stations went down too often, the data could be misrepresentative of true conditions throughout the year.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  4. A reduction turns a column into one number. Knowing how nulls are handled avoids silently biased statistics. NA means not available (missing).

    What NVIDIA says (1)

    “By default it’s value is set to True , we can change it to False to preserve NA values.”

    — cuDF: Working with missing data

Key terms: Null (missing value) Exploratory data analysis Descriptive statistics

Practice 4.1 (4 questions) Objective page

4.2 Visualization

Official objective: “Visualization”

Why visualize, Datashader for millions of points and interactive dashboards.

Key points

  1. Visualization means showing data as charts. The eye can spot unusual points and shapes quickly.

    What NVIDIA says (1)

    “Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns not easily surfaced by purely analytical methods.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  2. Overplotting happens when too many points overlap and hide the pattern. Datashader aggregates points into pixels to show density.

    What NVIDIA says (2)

    “The Datashader library directly supports cuDF and can rapidly render over millions of aggregated points.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “Datapoint rendering displaying high-resolution patterns is precisely what Datashader is designed for.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. Dash builds interactive web dashboards in Python. With cuDF behind it, the app can stay fast on large data.

    What NVIDIA says (1)

    “Plotly Dash enables data scientists to recast complex data and machine learning workflows as more accessible web applications.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  4. Precomputed aggregations are summaries prepared in advance. GPU speed lets the dashboard compute them on the fly as users interact.

    What NVIDIA says (1)

    “The use of Plotly’s Dash, RAPIDS, and Data shader allows users to build viz dashboards that both render datasets of 300 million+ rows and remain highly interactive without the need for precomputed aggregations.”

    — Making a Plotly Dash Census Viz Powered by RAPIDS

Key terms: Datashader Plotly Dash

Practice 4.2 (4 questions) Objective page

4.3 Choosing the right plot

Official objective: “Selecting appropriate plots for different analysis goals”

Histograms, heat maps, map plots and cross-filtering.

Key points

  1. A histogram counts how many values fall into each range. It shows the shape of a distribution, such as most trips being short.

    What NVIDIA says (1)

    “An hvPlot histogram of trip durations generated with the Divvy dataset In this instance, the vast majority of bike trips appear under 20 minutes.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  2. A heat map colors a grid by a value across two categories. It makes daily and weekly patterns easy to see.

    What NVIDIA says (1)

    “An hvPlot heat map showing trips by hour and day of week, per month Adding a widget for interactivity enables scrubbing through the months to search for patterns over a full year (Figure 2).”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. A hexbin chart groups nearby points into hexagon cells and colors them by count. It keeps maps readable for many points.

    What NVIDIA says (1)

    “Figure 3 shows the hexbin chart that aggregates trip start and ending locations to a manageable amount, verifying that the data is accurate to the bike share system map.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  4. Cross-filtering links charts so a selection in one filters the others. It replaces writing many separate groupby and query calls.

    What NVIDIA says (1)

    “Instead of creating several individual group by and query operations, a cuxfilter dashboard can simply cross-link numerous charts to quickly find patterns or anomalies (Figure 5).”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: cuxfilter Histogram Heat map Cross-filtering

Practice 4.3 (4 questions) Objective page

4.4 Hypothesis testing and significance

Official objective: “Hypothesis testing and statistical significance evaluation”

p-values, t-tests, confidence intervals and sample size.

Key points

  1. A hypothesis test asks whether an observed difference could easily happen by chance. The p-value is that chance under the 'no difference' assumption; a small p-value means significance.

    What NVIDIA says (1)

    “The biggest result is that, across all attempts, both of the lower-latency conditions (25 ms and 55 ms) improved the number of targets eliminated (Figure 3), a difference that was found to be statistically significant in pairwise t-tests ( p-value << 0.001).”

    — Improving Player Performance with Low Latency as Evident from FPS Aim Trainer Experiments

  2. A confidence interval is a range that likely contains the true value. More trials shrink it; too few trials leave comparisons inconclusive.

    What NVIDIA says (2)

    “A single success rate on N rollouts tells you almost nothing about how confident you should be in a policy’s true performance.”

    — How to Evaluate General-Purpose Robot Policies for Real-World Deployment

    “Most published benchmarks do not run a sufficient number of rollouts to achieve statistical significance when comparing the performance of two policies.”

    — How to Evaluate General-Purpose Robot Policies for Real-World Deployment

  3. The width of a confidence interval shrinks roughly with the square root of the sample size. So precision gets expensive.

    What NVIDIA says (1)

    “Narrowing the confidence interval from 10 to 2 percentage points requires roughly 15x more rollouts (70 to 1,030).”

    — How to Evaluate General-Purpose Robot Policies for Real-World Deployment

Key terms: p-value Confidence interval

Practice 4.4 (3 questions) Objective page

4.5 Patterns, trends and relationships

Official objective: “Interpreting patterns, trends, and relationships in data”

Pearson and Spearman correlation, rolling means and changing trends.

Key points

  1. Correlation measures how two variables move together. Causation needs more evidence, such as a controlled experiment.

    What NVIDIA says (2)

    “pearson : Standard correlation coefficient spearman : Spearman rank correlation”

    — cuDF API: DataFrame.corr

    “However, it’s important to remember that correlation and causation are two different things.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

  2. Monotonic means one variable tends to rise when the other rises, not always at the same rate. Spearman uses ranks, so it captures that without assuming a line.

    What NVIDIA says (1)

    “Method used to compute correlation: pearson : Standard correlation coefficient spearman : Spearman rank correlation”

    — cuDF API: DataFrame.corr

  3. A trend is the long-run direction of a series. A rolling window averages nearby values to smooth out noise.

    What NVIDIA says (1)

    “Parameters : window int, offset or a BaseIndexer subclass Size of the window, i.e., the number of observations used to calculate the statistic.”

    — cuDF API: DataFrame.rolling

  4. A linear model assumes a straight-line relationship. When the pattern curves, a nonlinear model can describe it better.

    What NVIDIA says (2)

    “While there is a strong relationship between population and time, the relationship is not linear because various factors influence changes from year to year.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

    “Nonlinear regression can estimate models with arbitrary relationships between independent and dependent variables.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

Key terms: Correlation Rolling mean

Practice 4.5 (4 questions) Objective page