Data Analysis and Visualization

10% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI

5.1 Insights from large datasets

Official objective: “Awareness of the process of extracting insights from large datasets using data mining, data visualization, and similar techniques.”

Scaling exploration with cuDF, finding gaps, clustering and early visualization.

Key points

  1. pandas runs on one CPU core and slows past 1-2 GB. cuDF parallelizes the same kind of operations across GPU cores.

    What NVIDIA says (2)

    “pandas is designed to run on a single core and starts slowing down when data size hits 1-2 GB”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

    “For that middle ground of 2-10 GB, RAPIDS cuDF is the Goldilocks solution that is just right.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  2. Exploratory data analysis (EDA) is the first look at data to learn its shape and problems. Gaps affect whether the data can be relied on alone.

    What NVIDIA says (1)

    “However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  3. Clustering groups similar items without labels. It can reveal unknown patterns, such as customer segments or image themes.

    What NVIDIA says (1)

    “Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”

    — What is Machine Learning and Why Does It Matter?

  4. Anscombe's quartet shows datasets with the same statistics but very different shapes. Looking early catches such surprises.

    What NVIDIA says (1)

    “Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: pandas RAPIDS cuDF Exploratory data analysis Clustering

Practice 5.1 (4 questions)

5.2 Attention maps

Official objective: “Develop content for attention maps in multimodal settings.”

What attention computes, how cross-attention keys shape attention maps, and attention as explanation.

Key points

  1. Attention lets a model weigh other parts of the input when processing one part. An attention map is a picture of those weights, for example over image regions or words.

    What NVIDIA says (1)

    “Attention units follow these tags, calculating a kind of algebraic map of how each element relates to the others.”

    — What Is a Transformer Model?

  2. Cross-attention lets image features attend to prompt words. Each word gets an attention map over the image. NVIDIA found the keys decide where those maps fall. VAE means variational autoencoder.

    What NVIDIA says (1)

    “Our main insight is that the key pathway of the cross-attention module in the diffusion model (the K matrix) controls the layout of the attention maps.”

    — NVIDIA Technical Blog: Personalizing Text-to-Image Models

  3. Reading attention maps helps diagnose models. Leaking attention means the new word affects parts of the image it should not.

    What NVIDIA says (1)

    “existing techniques tend to overfit that component, causing the attention on”

    — NVIDIA Technical Blog: Personalizing Text-to-Image Models

  4. Attention maps highlight which input regions or tokens mattered most. That makes a single decision easier to inspect.

    What NVIDIA says (1)

    “forcing the model itself to show its work.”

    — What Is Explainable AI (XAI)?

Key terms: Cross-attention Attention map

Practice 5.2 (4 questions)

5.3 Graphs and charts

Official objective: “Create graphs, charts, or other visualizations to convey the results of data analysis using specialized software.”

Histograms, heat maps and cross-filtering dashboards.

Key points

  1. A histogram shows how often values fall in each range. NVIDIA's example used one to show most bike trips were under 20 minutes.

    What NVIDIA says (2)

    “An hvPlot histogram of trip durations generated with the Divvy dataset”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “In this instance, the vast majority of bike trips appear under 20 minutes.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  2. hvPlot is a pandas-like plotting API with built-in interactivity.

    What NVIDIA says (1)

    “Charts in hvPlot can be interactively displayed using Bokeh and Plotly extensions, or statically with the Matplotlib extension.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. Cross-filtering means selecting data in one chart filters all linked charts. NCCL means the NVIDIA Collective Communications Library.

    What NVIDIA says (1)

    “cuxfilter enables GPU accelerated cross-filtering dashboards from notebooks, in just a few lines of Python code.”

    — Welcome to cuxfilter’s documentation — cuxfilter 26.06.00 documentation

  4. A heat map colors a grid of two categories by a value. It shows patterns across both at once.

    What NVIDIA says (1)

    “An hvPlot heat map showing trips by hour and day of week, per month”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: Histogram Heat map Cross-filtering

Practice 5.3 (4 questions)

5.4 Relationships, trends and confounders

Official objective: “Identify relationships and trends or any factors that could affect the results of research.”

Correlation versus cause, site-specific confounders and spotting trends.

Key points

  1. Correlation means two variables move together. It is not proof of cause. A hidden third factor, a confounder, can drive both.

    What NVIDIA says (2)

    “For any use case, be aware of the variables that are highly correlated, as they could skew results.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

    “Take this into account if you are using any future algorithm that assumes the variables are independent, such as linear regression .”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  2. A confounding factor affects both inputs and outcomes, making a false link look real. Single-site data can carry such site-specific bias.

    What NVIDIA says (1)

    “Medical institutions have had to rely on their own data sources, which can be biased by, for example, patient demographics, the instruments used or clinical specializations.”

    — What Is Federated Learning?

  3. Linking charts lets you slice data quickly and spot trends over time.

    What NVIDIA says (1)

    “a clear pattern emerges between weekday and weekend trips”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: Cross-filtering Correlation Confounding factor

Practice 5.4 (3 questions)