2.3 Machine learning fundamentals
Feature engineering, model comparison with grid search and cross-validation, and supervised versus unsupervised learning.
Key points
Here a transformer is a scikit-learn preprocessing step, not a neural network. Feature engineering means shaping raw data into useful inputs, such as scaled numbers.
What NVIDIA says (2)
“Transformers apply algorithms to clean or reshape the training data before it is fed into a model.”
“For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”
A hyperparameter is a setting chosen before training, such as tree depth. Grid search tries many settings. Cross-validation scores each fairly on several splits.
What NVIDIA says (1)
“For effective model selection, scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”
A label is the known answer for a training example. Clustering is a common unsupervised task.
What NVIDIA says (1)
“Supervised learning algorithms use labeled data, unsupervised learning algorithms find patterns in unlabeled data.”
Key terms
- Cross-validation: Training and validating several times on different splits of the data to estimate accuracy fairly.
- scikit-learn: A Python library for traditional machine learning built on NumPy.
- Feature engineering: Shaping raw data into useful model inputs, for example by scaling or encoding.
- Hyperparameter: A setting chosen before training, such as learning rate or LoRA rank.
Sample question
In scikit-learn, what do transformers do before data reaches the model?
Show the answer
Answer: Clean or reshape the data, for example feature scaling or encoding categorical variables
Here a transformer is a scikit-learn preprocessing step, not a neural network. Feature engineering means shaping raw data into useful inputs, such as scaled numbers.
What NVIDIA says (2)
“Transformers apply algorithms to clean or reshape the training data before it is fed into a model.”
“For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”
Practice 2.3 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 2.2 Multimodal loss functions · 2.4 Nonsequential networks and residual connections →