2.2 Regression, classification and clustering
Picking a model family: linear and logistic regression, k-means, DBSCAN and random forests.
Key points
Regression predicts a continuous number. Classification predicts a category. Clustering groups unlabeled data. PCA means principal component analysis.
What NVIDIA says (1)
“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”
Logistic regression is a classification model that outputs a probability for each class. UMAP means Uniform Manifold Approximation and Projection.
What NVIDIA says (1)
“Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”
Unsupervised learning finds patterns in data without labels. Clustering is its most common task. DBSCAN means Density-Based Spatial Clustering of Applications with Noise.
What NVIDIA says (2)
“Unsupervised learning algorithms attempt to ‘learn’ patterns in unlabeled data sets, discovering similarities, or regularities.”
“Common unsupervised tasks include clustering and association.”
k-means assigns every point to the nearest of k centers. DBSCAN instead grows clusters from dense regions and can mark sparse points as noise. DBSCAN means Density-Based Spatial Clustering of Applications with Noise.
What NVIDIA says (2)
“DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”
“This also allows DBSCAN to be robust to noise.”
Overfitting means a model learns noise in the training data and does worse on new data. Random forests average many varied trees, which smooths out the noise.
What NVIDIA says (2)
“Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”
“Whereas decision trees are based upon a fixed set of features, and often overfit, randomness is critical to the success of the forest.”
Key terms
- Supervised learning: Learning from examples that come with the correct answer (labels).
- Unsupervised learning: Finding patterns in data that has no labels.
- Regression: Predicting a number, such as a price.
- Classification: Predicting a category, such as spam or not spam.
- Clustering: Grouping similar unlabeled rows together.
- DBSCAN: A clustering method that finds dense regions of any shape and marks sparse points as noise.
- Random forest: Many decision trees trained on random samples whose votes are combined.
Sample question
You must predict a house's sale price. Which kind of task is this, and which model fits as a simple baseline?
Show the answer
Answer: Regression; linear regression
Regression predicts a continuous number. Classification predicts a category. Clustering groups unlabeled data. PCA means principal component analysis.
What NVIDIA says (1)
“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”
Practice 2.2 (5 questions) Full Machine Learning With RAPIDS guide
← 2.1 Training on GPUs with cuML and XGBoost · 2.3 Evaluating and comparing models →