2.2 Regression, classification and clustering

NCA-ADS · Machine Learning With RAPIDS (16% of the exam) · Official objective: “Regression, classification, and clustering techniques”

Picking a model family: linear and logistic regression, k-means, DBSCAN and random forests.

Key points

  1. Regression predicts a continuous number. Classification predicts a category. Clustering groups unlabeled data. PCA means principal component analysis.

    What NVIDIA says (1)

    “Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

  2. Logistic regression is a classification model that outputs a probability for each class. UMAP means Uniform Manifold Approximation and Projection.

    What NVIDIA says (1)

    “Logistic regression is an algorithm used for classification to predict the probability that an item belongs to a class, for example the probability that an email is spam.”

    — NVIDIA Glossary: Linear Regression and Logistic Regression

  3. Unsupervised learning finds patterns in data without labels. Clustering is its most common task. DBSCAN means Density-Based Spatial Clustering of Applications with Noise.

    What NVIDIA says (2)

    “Unsupervised learning algorithms attempt to ‘learn’ patterns in unlabeled data sets, discovering similarities, or regularities.”

    — NVIDIA Glossary: K-Means

    “Common unsupervised tasks include clustering and association.”

    — NVIDIA Glossary: K-Means

  4. k-means assigns every point to the nearest of k centers. DBSCAN instead grows clusters from dense regions and can mark sparse points as noise. DBSCAN means Density-Based Spatial Clustering of Applications with Noise.

    What NVIDIA says (2)

    “DBSCAN is a very powerful yet fast clustering technique that finds clusters where data is concentrated.”

    — cuML API: DBSCAN

    “This also allows DBSCAN to be robust to noise.”

    — cuML API: DBSCAN

  5. Overfitting means a model learns noise in the training data and does worse on new data. Random forests average many varied trees, which smooths out the noise.

    What NVIDIA says (2)

    “Random forest uses a technique called “bagging” to build full decision trees in parallel from random bootstrap samples of the data set and features.”

    — NVIDIA Glossary: Random Forest

    “Whereas decision trees are based upon a fixed set of features, and often overfit, randomness is critical to the success of the forest.”

    — NVIDIA Glossary: Random Forest

Key terms

Sample question

You must predict a house's sale price. Which kind of task is this, and which model fits as a simple baseline?

Show the answer

Answer: Regression; linear regression

Regression predicts a continuous number. Classification predicts a category. Clustering groups unlabeled data. PCA means principal component analysis.

What NVIDIA says (1)

“Linear regression is an algorithm used for regression to predict a numeric value, for example the price of a house.”

— NVIDIA Glossary: Linear Regression and Logistic Regression

Practice 2.2 (5 questions) Full Machine Learning With RAPIDS guide

← 2.1 Training on GPUs with cuML and XGBoost · 2.3 Evaluating and comparing models →