Keyword Search and Inverted Indexes
Understand how keyword search really works, from inverted indexes to BM25, and tune it for your content.
A taste of a lesson
Why does a short document with my keyword once outrank a long one that mentions it ten times?
Two parts of BM25 cause that. First, term frequency saturates: the second and third mentions add less than the first, and the tenth adds very little. Second, length normalisation: a long document is expected to contain more words, so each match counts for less relative to the average document length. A short document where the term is a large share of the text looks more focused on it. Which behaviour would you prefer for your content, and why?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Build and explain an inverted index
- Choose text analysis steps such as stemming and stop words for your content
- Explain how TF-IDF and BM25 rank documents
- Use field boosts, phrase queries and fuzzy matching effectively
- Recognise vocabulary mismatch and reduce it with expansion or hybrid search
Lesson plan
- 1 The inverted index Build a small inverted index by hand and use it to answer queries. Start
- 2 Text analysis Decide how text is split and normalised before indexing. Start
- 3 Scoring with TF-IDF and BM25 Understand how keyword engines rank matching documents. Start
- 4 Query features Use phrases, boosts and fuzzy matching to improve results. Start
- 5 Vocabulary mismatch Reduce misses caused by different words for the same thing. Start
- 6 Languages and evaluation Handle other languages and measure the effect of changes. Start
Try asking
About this tutor
For developers who want to understand the search technique that still powers most search boxes and remains essential in hybrid RAG systems. You learn how an inverted index works, how text analysis (tokenising, lowercasing, stemming, stop words) shapes what matches, how TF-IDF and BM25 rank documents, and how to use field boosts, phrase queries and fuzzy matching. You also see where keyword search shines (codes, names, exact phrases) and fails (synonyms and paraphrase), and how query expansion and analysers for other languages help. Concepts apply to search engines and database full text search alike.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Search engineer teaching embeddings, chunking, vector and keyword search, and reranking from first principles
9 tutors 374 lessons taught Sample
I come from search: indexes, ranking and the long tail of queries that make a search box look foolish. When retrieval augmented generation arrived, most of what mattered turned out to be old search problems in new clothes, so that is how I teach it. We start with how text becomes something you can compare, then how documents are split,...
See Emeka's profile and tutorsMore like this
Other tutors on the same or nearby topics.