1.5 Class imbalance and synthetic data
Why accuracy misleads, class weights, stratified splits and synthetic data.
Key points
Class imbalance means one label is much rarer than another. A balanced class weight makes errors on the rare class cost more during training.
What NVIDIA says (2)
“The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”
“If 'balanced' , class weights are computed from the training labels.”
A stratified split keeps each class's share the same in every part. Without it, a small test set may contain few or no rare cases.
What NVIDIA says (1)
“If not None, data is split in a stratified fashion, using this as the class labels.”
Synthetic data is artificial data made by rules, simulations or generative models. It can fill gaps in rare classes and protect privacy.
What NVIDIA says (2)
“Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”
“Data Privacy : Synthetic data helps overcome privacy issues by generating training data that mimic real-world statistics without directly corresponding to individual records.”
Accuracy is the share of correct predictions. On imbalanced data it hides poor results on the rare class; a confusion matrix shows them.
What NVIDIA says (2)
“Accuracy classification score.”
“Compute confusion matrix to evaluate the accuracy of a classification.”
Key terms
- Class imbalance: When one class is much rarer than the others in the data.
- Stratified split: A train-test split that keeps the same class ratio in both parts.
- Synthetic data: Generated data that mimics real data, used to fill gaps or protect privacy.
- Accuracy: The share of predictions that are correct.
Try it
Sample question
Only 2% of transactions are fraud. Which cuML classifier setting gives the rare class more weight automatically?
Show the answer
Answer: class_weight='balanced', which weights classes inversely to their frequency
Class imbalance means one label is much rarer than another. A balanced class weight makes errors on the rare class cost more during training.
What NVIDIA says (2)
“The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”
“If 'balanced' , class weights are computed from the training labels.”
Practice 1.5 (4 questions) Full Data Manipulation and Preparation guide
← 1.4 Feature engineering for numbers and categories · 1.6 Dimensionality reduction and sampling →