Image embeddings and visual search
Build image similarity and text to image search, and measure whether results are actually relevant
A taste of a lesson
My product photo search returns items photographed on the same white table, not similar products. What is going on?
The embedding is picking up the whole image, and a shared background and lighting setup can dominate the vector, especially if many products were shot the same way. Three things to try: crop to the product (a detector or a simple centre crop) before embedding; compare a model trained more for object identity, such as a contrastive or self supervised one; and, if you have pairs of the same product in different photos, fine tune on them. Measure each change on a small judged query set. Quick check: how would you build ten test queries that would expose the background problem?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what image embeddings are and how their source model shapes similarity
- Compute cosine similarity and set a near duplicate threshold from labelled pairs
- Choose between exact and approximate nearest neighbour search
- Evaluate visual search with a judged query set and recall at k
- Identify bias and privacy risks, especially for images of people
Lesson plan
- 1 What an embedding captures Understand embeddings as vectors where distance reflects a learned notion of similarity. Start
- 2 Contrastive image and text models See how shared image and text spaces make search by description possible. Start
- 3 Similarity and near duplicates Compute cosine similarity and choose duplicate thresholds from evidence. Start
- 4 Indexing and approximate search Choose a search approach that fits collection size and speed needs. Start
- 5 Evaluating search quality Build a judged query set and measure recall at k and top result precision. Start
- 6 Domain tuning, bias and privacy Adapt embeddings to your data and handle people in images responsibly. Start
Try asking
About this tutor
For learners who want to find similar images, detect near duplicates or search photos with text. You will learn what an image embedding is, how embeddings from a classifier differ from those of contrastive image and text models, and how cosine similarity ranks results. Then you build up a search system in general terms: indexing vectors, approximate nearest neighbour search and its speed versus recall trade off, and thresholds for duplicate detection. A full lesson covers evaluation with recall at k and a small judged query set, and another covers adapting embeddings to your own domain. The course ends with bias and privacy, especially for images of people and faces.
Reviews
4.7
3 ratingsSample
- Kofi A.Sample
The privacy lesson was taken seriously, not tacked on. Helped me push back on a face matching feature request at work.
- Beatriz C.Sample
Clear on exact versus approximate search. I appreciated that it said a vector database was not needed for my small archive.
- Hiroshi N.Sample
Cropping before embedding fixed most of my background problem. The judged query set idea gave me a number to track instead of vibes.
About the teacher
Computer vision taught through real images, real failure cases and careful evaluation
9 tutors 406 lessons taught Sample
I teach computer vision: classification, detection, segmentation, document understanding, video and the newer models that combine images with language. Most of my work has been building vision systems that had to hold up outside the lab, under odd lighting, unusual cameras and labels that were not quite consistent. So my lessons spend as much time on data and evaluation as...
See Noor's profile and tutorsMore like this
Other tutors on the same or nearby topics.