1.6 Dimensionality reduction and sampling

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “Dimensionality reduction and data sampling”

PCA, UMAP, random projection and reproducible sampling.

Key points

  1. Dimensionality reduction means representing data with fewer columns while keeping most of the information. PCA keeps the directions with the most variance. PCA means principal component analysis.

    What NVIDIA says (1)

    “PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”

    — cuML API: PCA

  2. UMAP builds a nearest-neighbor graph and embeds the data into fewer dimensions. It keeps local neighborhoods, which makes clusters visible. UMAP means Uniform Manifold Approximation and Projection.

    What NVIDIA says (2)

    “UMAP is a popular dimension reduction algorithm used in fields like bioinformatics, NLP topic modeling, and ML preprocessing.”

    — Even Faster and More Scalable UMAP on the GPU with RAPIDS cuML

    “It works by creating a k-nearest neighbors (k-NN) graph, which is known in literature as an all-neighbors graph, to build a fuzzy topological representation of the data, which is used to embed high-dimensional data into lower dimensions.”

    — Even Faster and More Scalable UMAP on the GPU with RAPIDS cuML

  3. Sampling picks a subset of rows to work with. A fixed random_state gives the same sample every time.

    What NVIDIA says (2)

    “This function will always produce the same sample given an identical random_state .”

    — cuDF API: DataFrame.sample

    “frac float, optional Fraction of axis items to return.”

    — cuDF API: DataFrame.sample

  4. Random projection multiplies data by a random matrix to get fewer dimensions. The Johnson-Lindenstrauss lemma bounds how many dimensions keep distances roughly intact. PCA means principal component analysis.

    What NVIDIA says (2)

    “Reduce dimensionality through Gaussian random projection.”

    — cuML API: GaussianRandomProjection

    “n_components can be automatically adjusted according to the number of samples in the dataset and the bound given by the Johnson-Lindenstrauss lemma.”

    — cuML API: GaussianRandomProjection

Key terms

Sample question

What does PCA do?

Show the answer

Answer: It combines features into linear components, each capturing as much remaining variance as possible, to reduce dimensions

Dimensionality reduction means representing data with fewer columns while keeping most of the information. PCA keeps the directions with the most variance. PCA means principal component analysis.

What NVIDIA says (1)

“PCA (Principal Component Analysis) is a fundamental dimensionality reduction technique used to combine features in X in linear combinations such that each new component captures the most”

— cuML API: PCA

Practice 1.6 (4 questions) Full Data Manipulation and Preparation guide

← 1.5 Class imbalance and synthetic data · 1.7 Efficient storage with Parquet and modern frameworks →