1.4 Feature engineering for numbers and categories

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “Feature engineering for numerical and categorical variables”

One-hot and target encoding, standard scaling and binning.

Key points

  1. One-hot encoding turns a categorical column into one true/false column per category. Categorical means the values are labels, not numbers.

    What NVIDIA says (1)

    “cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”

    — cuDF API: get_dummies

  2. Target encoding replaces a category with a statistic of the label, usually its mean. cuML's TargetEncoder uses folds to avoid label leakage, which is when the label sneaks into a feature.

    What NVIDIA says (2)

    “The input data is grouped by the columns Xs and the aggregated mean value of Y of each group is calculated to replace each value of Xs .”

    — cuML API: TargetEncoder

    “Several optimizations are applied to prevent label leakage and parallelize the execution.”

    — cuML API: TargetEncoder

  3. Scaling puts numeric features on a similar range so no feature dominates by its units. StandardScaler computes z = (x - mean) / std and reuses those training statistics later.

    What NVIDIA says (2)

    “Standardize features by removing the mean and scaling to unit variance”

    — cuML API: StandardScaler

    “Mean and standard deviation are then stored to be used on later data using transform() .”

    — cuML API: StandardScaler

  4. Binning, also called discretization, groups continuous numbers into intervals. Each bin can then be treated as a category.

    What NVIDIA says (1)

    “Bin continuous data into intervals.”

    — cuML API: KBinsDiscretizer

  5. With all categories present, one column can be computed from the others. Dropping one removes that redundancy but changes how a penalized model treats categories.

    What NVIDIA says (1)

    “However, dropping one category breaks the symmetry of the original representation and can therefore induce a bias in downstream models, for instance for penalized linear classification or regression models.”

    — cuML API: OneHotEncoder

Key terms

Try it

Sample question

A 'color' column has values red, green and blue. Which cuDF call turns it into one indicator column per value?

Show the answer

Answer: cudf.get_dummies(df)

One-hot encoding turns a categorical column into one true/false column per category. Categorical means the values are labels, not numbers.

What NVIDIA says (1)

“cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”

— cuDF API: get_dummies

Practice 1.4 (5 questions) Full Data Manipulation and Preparation guide

← 1.3 GPU-accelerated ETL with RAPIDS, Dask and Spark · 1.5 Class imbalance and synthetic data →