Dimensionality Reduction in Practice
Use PCA, t-SNE and UMAP well, and avoid reading too much into pretty plots
A taste of a lesson
My UMAP plot of customer embeddings shows a small cluster far from the others. Is that group very different?
Maybe, but the plot alone cannot tell you. UMAP preserves neighbourhoods better than global distances, so 'far away' in the picture does not reliably mean 'very different' in the original space. Check three things: does the cluster persist across different neighbour settings and seeds; what do those customers have in common when you look at their actual features; and are they also separated when you cluster in the original embedding space? If all three agree, it is likely a real group. Which check can you run first?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Decide when dimensionality reduction will help a project
- Apply PCA as a preprocessing step and choose the number of components
- Produce t-SNE and UMAP visualisations with sensible settings
- Interpret neighbourhood plots without over reading distances or sizes
- Choose between feature selection and feature extraction
Lesson plan
- 1 Why reduce at all Identify goals that justify dimensionality reduction. Start
- 2 PCA as preprocessing Use PCA inside a modelling pipeline. Start
- 3 Visualising with t-SNE Visualise local structure with t-SNE and its settings. Start
- 4 Visualising with UMAP Use UMAP and understand how it differs from t-SNE. Start
- 5 Reading the plots Draw only conclusions that survive changes in settings. Start
- 6 Beyond the basics Know when autoencoders or selection methods fit better. Start
Try asking
About this tutor
An intermediate tutor for practitioners who want to reduce many features to fewer, for modelling, compression or visualisation. You will decide when reduction helps, apply PCA as a preprocessing step, and use t-SNE and UMAP to visualise embeddings and high dimensional data. A large part of the course is about interpretation: what distances, cluster sizes and gaps in these plots do and do not mean, how hyperparameters change the picture, and why clustering on a visualisation can mislead. You also compare feature selection with feature extraction, with notes on autoencoders.
Reviews
4.3
3 ratingsSample
- Ilse V.Sample
Stopped me from presenting t-SNE cluster sizes as meaningful. Running several seeds is now part of my routine.
- Yan L.Sample
Clear explanation of the difference between t-SNE and UMAP settings. Helpful for visualising our document embeddings.
- Kofi A.Sample
Good practical guidance on choosing PCA components through validation. The autoencoder lesson was quite brief.
About the teacher
Linear algebra for AI, with geometry first and notation second
9 tutors 338 lessons taught Sample
I teach the linear algebra behind modern AI: vectors, matrices, similarity, eigenvectors and the methods built on them, such as PCA, clustering and recommender systems. I trained in applied mathematics and later worked on search and recommendation features, so I like to connect each idea to something a real system does. My lessons begin with pictures and small numbers you...
See Katarzyna's profile and tutorsMore like this
Other tutors on the same or nearby topics.