1.4 Feature engineering for numbers and categories
One-hot and target encoding, standard scaling and binning.
Key points
One-hot encoding turns a categorical column into one true/false column per category. Categorical means the values are labels, not numbers.
What NVIDIA says (1)
“cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”
Target encoding replaces a category with a statistic of the label, usually its mean. cuML's TargetEncoder uses folds to avoid label leakage, which is when the label sneaks into a feature.
What NVIDIA says (2)
“The input data is grouped by the columns Xs and the aggregated mean value of Y of each group is calculated to replace each value of Xs .”
“Several optimizations are applied to prevent label leakage and parallelize the execution.”
Scaling puts numeric features on a similar range so no feature dominates by its units. StandardScaler computes z = (x - mean) / std and reuses those training statistics later.
What NVIDIA says (2)
“Standardize features by removing the mean and scaling to unit variance”
“Mean and standard deviation are then stored to be used on later data using transform() .”
Binning, also called discretization, groups continuous numbers into intervals. Each bin can then be treated as a category.
What NVIDIA says (1)
“Bin continuous data into intervals.”
With all categories present, one column can be computed from the others. Dropping one removes that redundancy but changes how a penalized model treats categories.
What NVIDIA says (1)
“However, dropping one category breaks the symmetry of the original representation and can therefore induce a bias in downstream models, for instance for penalized linear classification or regression models.”
Key terms
- Feature: An input column a model learns from.
- Feature engineering: Creating or reshaping input columns so a model can learn better.
- One-hot encoding: Turning a category column into one 0/1 column per category.
- Target encoding: Replacing each category with the average target value for that category.
- Standard scaling: Rescaling a numeric feature to mean 0 and standard deviation 1.
- Binning: Grouping a continuous number into ranges, such as age bands.
Try it
Sample question
A 'color' column has values red, green and blue. Which cuDF call turns it into one indicator column per value?
Show the answer
Answer: cudf.get_dummies(df)
One-hot encoding turns a categorical column into one true/false column per category. Categorical means the values are labels, not numbers.
What NVIDIA says (1)
“cudf . get_dummies ( df ) b a_value1 a_value2 0 0 True False 1 0 False True 2 0 False False”
Practice 1.4 (5 questions) Full Data Manipulation and Preparation guide
← 1.3 GPU-accelerated ETL with RAPIDS, Dask and Spark · 1.5 Class imbalance and synthetic data →