Data Analysis and Visualization
14% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI
4.1 Insights from large datasets
Interactive, GPU-accelerated exploration.
Key points
Visualization helps the brain take in large amounts of information quickly. NVIDIA notes interactions slower than 7-10 seconds disrupt the user's short-term memory. Graphics processing unit (GPU) libraries bring compute and render times down to interactive speeds.
What NVIDIA says (3)
“However, interactions such as filtering, selecting, or rerendering points that are slower than 7-10 seconds result in a disruption of a user”
“This style of visualization is essentially a hack for the brain to understand large amounts of information quickly.”
“Visualization compute and render times are brought down to interactive speeds”
NVIDIA's guide says that when an exploratory data analysis (EDA) workflow processes more than 2 gigabytes (GB) with compute-intensive tasks, central processing unit (CPU)-based tools can slow iteration. Switching to pandas-like RAPIDS graphics processing unit (GPU) libraries such as cuDF keeps the pace up.
What NVIDIA says (2)
“But when an EDA workflow is processing data larger than 2 GB, and requires compute intensive tasks, CPU-based solutions can start to constrain the iterative exploration process.”
“Replacing CPU-based libraries with the pandas-like RAPIDS GPU-accelerated libraries (such as cuDF) means you can keep a swift pace for your EDA process”
cuDF is a pandas-like RAPIDS graphics processing unit (GPU)-accelerated library. NVIDIA says it gives GPU acceleration through a familiar pandas-like application programming interface (API).
What NVIDIA says (2)
“enables GPU acceleration that unlocks access to your data insights through a familiar pandas-like API.”
“the pandas-like RAPIDS GPU-accelerated libraries (such as cuDF)”
Key terms: cuDF Exploratory data analysis
4.2 Comparing models with statistics
Regression metrics and what they do and do not tell you.
Key points
R² (the coefficient of determination) is the proportion of variance in the target that the model explains. A higher value means a better fit when you compare models trained on the same dataset.
What NVIDIA says (3)
“also known as the coefficient of determination, represents the proportion of variance explained by a model.”
“R² is a relative metric; that is, it can be used to compare with other models trained on the same dataset. A higher value indicates a better fit.”
“corresponds to the degree to which the variance in the dependent variable (the target) can be explained by the”
A residual is the difference between the actual and the predicted value. Mean squared error (MSE) is the average of the squared residuals. Squaring puts a much heavier penalty on large errors, so MSE is not robust to outliers.
What NVIDIA says (3)
“As the residuals are squared, MSE puts a significantly heavier penalty on large errors. Some of those might be outliers, so MSE is not robust to their presence.”
“a residual is a difference between the actual value and the predicted value.”
“The difference is that you are now interested in the average error instead of the total error.”
MAE (mean absolute error) uses absolute values, so it ignores the direction of errors. Like mean squared error (MSE) and RMSE, it is scale-dependent, so you cannot compare it between different datasets.
What NVIDIA says (2)
“Similar to MSE and RMSE, MAE is also scale-dependent, so you cannot compare it between different datasets.”
“Absolute value disregards the direction of the errors”
RMSE is the square root of mean squared error (MSE). Taking the root brings the metric back to the scale of the target variable, so it is easier to interpret.
What NVIDIA says (2)
“(RMSE) is closely related to MSE, as it is simply the square root of the latter.”
“Take the square to bring the metric back to the scale of the target variable, so it is easier to interpret and understand.”
Key terms: Residual R² MSE / RMSE / MAE
Try it: Metrics lab
4.3 Doing the data analysis
A practical order of work for a new dataset.
Key points
NVIDIA's exploratory data analysis (EDA) tutorial starts by reviewing the dataset and understanding the variables. This tells you the dimensions and the kinds of data in the DataFrame.
What NVIDIA says (2)
“Review the dataset and understand the variables that you are working with.”
“This helps you understand the dimensions of and the kinds of data in the DataFrame.”
NVIDIA's exploratory data analysis (EDA) walkthrough notes that some missing data is acceptable, but frequent outages could make the data misrepresent true conditions.
What NVIDIA says (1)
“Some missing data is acceptable, but if the stations went down too often, the data could be misrepresentative of true conditions throughout the year.”
NVIDIA's pandas glossary says pandas imports and exports comma-separated values (CSV), Structured Query Language (SQL) and spreadsheet files. Combined with its manipulation features, it can clean, shape and analyze tabular data.
What NVIDIA says (2)
“pandas facilitates importing and exporting datasets from various file formats, such as CSV, SQL, and spreadsheets.”
“These operations, combined with its data manipulation capabilities, enable pandas to clean, shape, and analyze tabular and statistical data.”
Key terms: pandas / DataFrame Exploratory data analysis
4.4 Charts and dashboards
Which visual and which tool for which question.
Key points
cuxfilter is a RAPIDS library for dashboards. NVIDIA's docs say it creates graphics processing unit (GPU)-accelerated cross-filtering dashboards from notebooks in a few lines of Python. Cross-filtering replaces hand-written DataFrame queries with a graphical user interface (GUI).
What NVIDIA says (2)
“cuxfilter enables GPU accelerated cross-filtering dashboards from notebooks, in just a few lines of Python code.”
“This approach replaces dataframe queries with a GUI tool.”
NVIDIA's visualization guide says hvPlot charts can be shown interactively using Bokeh and Plotly extensions, or statically with the Matplotlib extension.
What NVIDIA says (1)
“Charts in hvPlot can be interactively displayed using Bokeh and Plotly extensions, or statically with the Matplotlib extension.”
NVIDIA's visualization guide plots an hvPlot histogram of trip durations. It shows that most bike trips are under 20 minutes.
What NVIDIA says (2)
“In this instance, the vast majority of bike trips appear under 20 minutes.”
“An hvPlot histogram of trip durations generated with the Divvy dataset”
Key terms: cuxfilter
4.5 Relationships, trends and confounders
Spotting patterns and the factors that can distort results.
Key points
NVIDIA warns that R² does not measure bias, so an overfitted model can have a high R², and you should also look at other metrics.
What NVIDIA says (1)
“Second, R² does not give any measure of bias, so you can have an overfitted (highly biased) model with a high value of R².”
NVIDIA notes that data from a single institution can be biased by patient demographics, instruments or clinical specializations, which can affect the results of a model trained on it.
What NVIDIA says (1)
“Medical institutions have had to rely on their own data sources, which can be biased by, for example, patient demographics, the instruments used or clinical specializations.”
Cross-filtering replaces hand-written DataFrame queries with a graphical user interface (GUI). In NVIDIA's example, a clear pattern emerged between weekday and weekend trips.
What NVIDIA says (2)
“As shown in Figure 5, a clear pattern emerges between weekday and weekend trips”
“This approach replaces dataframe queries with a GUI tool.”