3.4 Augmenting and integrating datasets
Synthetic data for rare cases, derived features and safe joins.
Key points
Data augmentation adds new training examples based on existing data or simulations.
What NVIDIA says (2)
“Synthetically generated data can augment existing data for larger, more representative datasets.”
“If data used to train self-driving cars underrepresents uncommon scenes such as extreme weather conditions or traffic accidents, synthetic data can help augment the diversity of these datasets to better”
Integration means combining data or results from several sources into one dataset. Adding a computed column is a simple form of it. EDA means exploratory data analysis.
What NVIDIA says (1)
“Augmenting the data using RAPIDS cuSpatial to quickly calculate distances also shows that most trips are relatively short.”
Integrating data with a merge can add rows when keys repeat and lose rows when keys are missing. In cuDF, a left join keeps every row of the left table.
What NVIDIA says (1)
“df_merged = df_a . merge ( df_b , on = [ 'key' ], how = 'left' )”
Key terms
- Join (merge): Combining two tables by matching values in a key column.
- Synthetic data: Generated data that mimics real data, used to fill gaps or protect privacy.
- Data augmentation: Adding new rows or derived columns to make training data richer.
Sample question
Rare but important cases are underrepresented in your training data. What can help?
Show the answer
Answer: Augment the dataset with synthetic data to make it larger and more representative
Data augmentation adds new training examples based on existing data or simulations.
What NVIDIA says (2)
“Synthetically generated data can augment existing data for larger, more representative datasets.”
“If data used to train self-driving cars underrepresents uncommon scenes such as extreme weather conditions or traffic accidents, synthetic data can help augment the diversity of these datasets to better”
Practice 3.4 (3 questions) Full Data Science Pipelines and Workflow Automation guide
← 3.3 Fixing underfitting and overfitting · 3.5 Automating and scaling workflows →