1.5 Class imbalance and synthetic data

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “Handling class imbalance and generating synthetic data”

Why accuracy misleads, class weights, stratified splits and synthetic data.

Key points

  1. Class imbalance means one label is much rarer than another. A balanced class weight makes errors on the rare class cost more during training.

    What NVIDIA says (2)

    “The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”

    — cuML API: LogisticRegression

    “If 'balanced' , class weights are computed from the training labels.”

    — cuML API: RandomForestClassifier

  2. A stratified split keeps each class's share the same in every part. Without it, a small test set may contain few or no rare cases.

    What NVIDIA says (1)

    “If not None, data is split in a stratified fashion, using this as the class labels.”

    — cuML API: train_test_split

  3. Synthetic data is artificial data made by rules, simulations or generative models. It can fill gaps in rare classes and protect privacy.

    What NVIDIA says (2)

    “Data Quality : Real-world datasets can be imbalanced, which can result in biased outputs from generative models and ML models.”

    — NVIDIA Glossary: Synthetic Data Generation

    “Data Privacy : Synthetic data helps overcome privacy issues by generating training data that mimic real-world statistics without directly corresponding to individual records.”

    — NVIDIA Glossary: Synthetic Data Generation

  4. Accuracy is the share of correct predictions. On imbalanced data it hides poor results on the rare class; a confusion matrix shows them.

    What NVIDIA says (2)

    “Accuracy classification score.”

    — cuML API: accuracy_score

    “Compute confusion matrix to evaluate the accuracy of a classification.”

    — cuML API: confusion_matrix

Key terms

Try it

Sample question

Only 2% of transactions are fraud. Which cuML classifier setting gives the rare class more weight automatically?

Show the answer

Answer: class_weight='balanced', which weights classes inversely to their frequency

Class imbalance means one label is much rarer than another. A balanced class weight makes errors on the rare class cost more during training.

What NVIDIA says (2)

“The “balanced” mode uses the values of y to automatically adjust weights inversely proportional to class frequencies in the input data as n_samples / (n_classes * np.bincount(y)) .”

— cuML API: LogisticRegression

“If 'balanced' , class weights are computed from the training labels.”

— cuML API: RandomForestClassifier

Practice 1.5 (4 questions) Full Data Manipulation and Preparation guide

← 1.4 Feature engineering for numbers and categories · 1.6 Dimensionality reduction and sampling →