3.4 Augmenting and integrating datasets

NCA-ADS · Data Science Pipelines and Workflow Automation (13% of the exam) · Official objective: “Dataset augmentation and integration for enhanced training data”

Synthetic data for rare cases, derived features and safe joins.

Key points

  1. Data augmentation adds new training examples based on existing data or simulations.

    What NVIDIA says (2)

    “Synthetically generated data can augment existing data for larger, more representative datasets.”

    — NVIDIA Glossary: Synthetic Data Generation

    “If data used to train self-driving cars underrepresents uncommon scenes such as extreme weather conditions or traffic accidents, synthetic data can help augment the diversity of these datasets to better”

    — What Is Trustworthy AI?

  2. Integration means combining data or results from several sources into one dataset. Adding a computed column is a simple form of it. EDA means exploratory data analysis.

    What NVIDIA says (1)

    “Augmenting the data using RAPIDS cuSpatial to quickly calculate distances also shows that most trips are relatively short.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

  3. Integrating data with a merge can add rows when keys repeat and lose rows when keys are missing. In cuDF, a left join keeps every row of the left table.

    What NVIDIA says (1)

    “df_merged = df_a . merge ( df_b , on = [ 'key' ], how = 'left' )”

    — cuDF API: DataFrame.merge

Key terms

Sample question

Rare but important cases are underrepresented in your training data. What can help?

Show the answer

Answer: Augment the dataset with synthetic data to make it larger and more representative

Data augmentation adds new training examples based on existing data or simulations.

What NVIDIA says (2)

“Synthetically generated data can augment existing data for larger, more representative datasets.”

— NVIDIA Glossary: Synthetic Data Generation

“If data used to train self-driving cars underrepresents uncommon scenes such as extreme weather conditions or traffic accidents, synthetic data can help augment the diversity of these datasets to better”

— What Is Trustworthy AI?

Practice 3.4 (3 questions) Full Data Science Pipelines and Workflow Automation guide

← 3.3 Fixing underfitting and overfitting · 3.5 Automating and scaling workflows →