Loading…
Reduce high-dimensional data for visualization, noise removal, and feature extraction. From the classic PCA to modern manifold techniques like UMAP.
8
Algorithms
2
Categories
2–10
Components
Project data onto linear subspaces that maximise variance or discover latent factors.
The gold standard for variance-based linear dimensionality reduction. Projects data onto orthogonal principal components that capture the most variance.
Key Features
Best for: General-purpose reduction, noise filtering, feature decorrelation
Process large datasets that don't fit in memory by computing PCA in mini-batches. Same result as standard PCA but with constant memory usage.
Key Features
Best for: Very large datasets (100K+ rows) where standard PCA runs out of memory
Identify hidden "latent" variables that explain observed correlations. Unlike PCA, it models measurement error and can reveal interpretable factors.
Key Features
Best for: Psychology surveys, test score analysis, discovering hidden constructs
Singular Value Decomposition without centring data first — ideal for sparse matrices like TF-IDF text features or one-hot encodings.
Key Features
Best for: Text data (TF-IDF), sparse matrices, topic modelling (LSA)
Preserve complex, curved structure that linear methods miss — ideal for visualization.
The most popular tool for high-quality 2D and 3D visual maps of high-dimensional data. Optimises for preserving local neighbourhood structure.
Key Features
Best for: Visualizing clusters in image embeddings, gene expression, word vectors
A modern, faster alternative to t-SNE that preserves both local and global structure. Scales better to large datasets and higher dimensions.
Key Features
Best for: Large-scale visual exploration, pre-processing step for clustering
Geometric reduction that preserves geodesic (shortest-path) distances along the data manifold, unlike Euclidean methods. Unfolds curved manifolds.
Key Features
Best for: Swiss-roll, face pose estimation, any data on a curved manifold
Uses the eigenvectors of the data's graph Laplacian to find low-dimensional structure. Particularly effective when data has clear cluster boundaries.
Key Features
Best for: Cluster separation, image segmentation, network/graph data