Data Analysis and Visualization

14% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI

4.1 Insights from large datasets

Official objective: “Awareness of the process of extracting insights from large datasets using data mining, data visualization, and similar techniques.”

Interactive, GPU-accelerated exploration.

Key points

  1. Visualization helps the brain take in large amounts of information quickly. NVIDIA notes interactions slower than 7-10 seconds disrupt the user's short-term memory. Graphics processing unit (GPU) libraries bring compute and render times down to interactive speeds.

    What NVIDIA says (3)

    “However, interactions such as filtering, selecting, or rerendering points that are slower than 7-10 seconds result in a disruption of a user”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “This style of visualization is essentially a hack for the brain to understand large amounts of information quickly.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “Visualization compute and render times are brought down to interactive speeds”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  2. NVIDIA's guide says that when an exploratory data analysis (EDA) workflow processes more than 2 gigabytes (GB) with compute-intensive tasks, central processing unit (CPU)-based tools can slow iteration. Switching to pandas-like RAPIDS graphics processing unit (GPU) libraries such as cuDF keeps the pace up.

    What NVIDIA says (2)

    “But when an EDA workflow is processing data larger than 2 GB, and requires compute intensive tasks, CPU-based solutions can start to constrain the iterative exploration process.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “Replacing CPU-based libraries with the pandas-like RAPIDS GPU-accelerated libraries (such as cuDF) means you can keep a swift pace for your EDA process”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. cuDF is a pandas-like RAPIDS graphics processing unit (GPU)-accelerated library. NVIDIA says it gives GPU acceleration through a familiar pandas-like application programming interface (API).

    What NVIDIA says (2)

    “enables GPU acceleration that unlocks access to your data insights through a familiar pandas-like API.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “the pandas-like RAPIDS GPU-accelerated libraries (such as cuDF)”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: cuDF Exploratory data analysis

Practice 4.1 (3 questions)

4.2 Comparing models with statistics

Official objective: “Compare models using statistical performance metrics, such as loss functions or proportion of explained variance.”

Regression metrics and what they do and do not tell you.

Key points

  1. R² (the coefficient of determination) is the proportion of variance in the target that the model explains. A higher value means a better fit when you compare models trained on the same dataset.

    What NVIDIA says (3)

    “also known as the coefficient of determination, represents the proportion of variance explained by a model.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “R² is a relative metric; that is, it can be used to compare with other models trained on the same dataset. A higher value indicates a better fit.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “corresponds to the degree to which the variance in the dependent variable (the target) can be explained by the”

    — A Comprehensive Overview of Regression Evaluation Metrics

  2. A residual is the difference between the actual and the predicted value. Mean squared error (MSE) is the average of the squared residuals. Squaring puts a much heavier penalty on large errors, so MSE is not robust to outliers.

    What NVIDIA says (3)

    “As the residuals are squared, MSE puts a significantly heavier penalty on large errors. Some of those might be outliers, so MSE is not robust to their presence.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “a residual is a difference between the actual value and the predicted value.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “The difference is that you are now interested in the average error instead of the total error.”

    — A Comprehensive Overview of Regression Evaluation Metrics

  3. MAE (mean absolute error) uses absolute values, so it ignores the direction of errors. Like mean squared error (MSE) and RMSE, it is scale-dependent, so you cannot compare it between different datasets.

    What NVIDIA says (2)

    “Similar to MSE and RMSE, MAE is also scale-dependent, so you cannot compare it between different datasets.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “Absolute value disregards the direction of the errors”

    — A Comprehensive Overview of Regression Evaluation Metrics

  4. RMSE is the square root of mean squared error (MSE). Taking the root brings the metric back to the scale of the target variable, so it is easier to interpret.

    What NVIDIA says (2)

    “(RMSE) is closely related to MSE, as it is simply the square root of the latter.”

    — A Comprehensive Overview of Regression Evaluation Metrics

    “Take the square to bring the metric back to the scale of the target variable, so it is easier to interpret and understand.”

    — A Comprehensive Overview of Regression Evaluation Metrics

Key terms: Residual R² MSE / RMSE / MAE

Try it: Metrics lab

Practice 4.2 (4 questions)

4.3 Doing the data analysis

Official objective: “Conduct data analysis under the supervision of a senior team member.”

A practical order of work for a new dataset.

Key points

  1. NVIDIA's exploratory data analysis (EDA) tutorial starts by reviewing the dataset and understanding the variables. This tells you the dimensions and the kinds of data in the DataFrame.

    What NVIDIA says (2)

    “Review the dataset and understand the variables that you are working with.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

    “This helps you understand the dimensions of and the kinds of data in the DataFrame.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  2. NVIDIA's exploratory data analysis (EDA) walkthrough notes that some missing data is acceptable, but frequent outages could make the data misrepresent true conditions.

    What NVIDIA says (1)

    “Some missing data is acceptable, but if the stations went down too often, the data could be misrepresentative of true conditions throughout the year.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  3. NVIDIA's pandas glossary says pandas imports and exports comma-separated values (CSV), Structured Query Language (SQL) and spreadsheet files. Combined with its manipulation features, it can clean, shape and analyze tabular data.

    What NVIDIA says (2)

    “pandas facilitates importing and exporting datasets from various file formats, such as CSV, SQL, and spreadsheets.”

    — What Is Pandas and Why Does it Matter?

    “These operations, combined with its data manipulation capabilities, enable pandas to clean, shape, and analyze tabular and statistical data.”

    — What Is Pandas and Why Does it Matter?

Key terms: pandas / DataFrame Exploratory data analysis

Practice 4.3 (3 questions)

4.4 Charts and dashboards

Official objective: “Create graphs, charts, or other visualizations to convey the results of data analysis using specialized software.”

Which visual and which tool for which question.

Key points

  1. cuxfilter is a RAPIDS library for dashboards. NVIDIA's docs say it creates graphics processing unit (GPU)-accelerated cross-filtering dashboards from notebooks in a few lines of Python. Cross-filtering replaces hand-written DataFrame queries with a graphical user interface (GUI).

    What NVIDIA says (2)

    “cuxfilter enables GPU accelerated cross-filtering dashboards from notebooks, in just a few lines of Python code.”

    — Welcome to cuxfilter’s documentation — cuxfilter 26.06.00 documentation

    “This approach replaces dataframe queries with a GUI tool.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  2. NVIDIA's visualization guide says hvPlot charts can be shown interactively using Bokeh and Plotly extensions, or statically with the Matplotlib extension.

    What NVIDIA says (1)

    “Charts in hvPlot can be interactively displayed using Bokeh and Plotly extensions, or statically with the Matplotlib extension.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. NVIDIA's visualization guide plots an hvPlot histogram of trip durations. It shows that most bike trips are under 20 minutes.

    What NVIDIA says (2)

    “In this instance, the vast majority of bike trips appear under 20 minutes.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “An hvPlot histogram of trip durations generated with the Divvy dataset”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: cuxfilter

Practice 4.4 (3 questions)

4.5 Relationships, trends and confounders

Official objective: “Identify relationships and trends or any factors that could affect the results of research.”

Spotting patterns and the factors that can distort results.

Key points

  1. NVIDIA warns that R² does not measure bias, so an overfitted model can have a high R², and you should also look at other metrics.

    What NVIDIA says (1)

    “Second, R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”

    — A Comprehensive Overview of Regression Evaluation Metrics

  2. NVIDIA notes that data from a single institution can be biased by patient demographics, instruments or clinical specializations, which can affect the results of a model trained on it.

    What NVIDIA says (1)

    “Medical institutions have had to rely on their own data sources, which can be biased by, for example, patient demographics, the instruments used or clinical specializations.”

    — What Is Federated Learning?

  3. Cross-filtering replaces hand-written DataFrame queries with a graphical user interface (GUI). In NVIDIA's example, a clear pattern emerged between weekday and weekend trips.

    What NVIDIA says (2)

    “As shown in Figure 5, a clear pattern emerges between weekday and weekend trips”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “This approach replaces dataframe queries with a GUI tool.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Practice 4.5 (3 questions)