Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Keyword Search and Inverted Indexes

Understand how keyword search really works, from inverted indexes to BM25, and tune it for your content.

Your first 3 replies from the tutor are free. Then you can continue the lesson for $6 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Keyword Search and Inverted Indexes AI tutor following Emeka Nwosu's plan
Student:

Why does a short document with my keyword once outrank a long one that mentions it ten times?

Tutor:

Two parts of BM25 cause that. First, term frequency saturates: the second and third mentions add less than the first, and the tenth adds very little. Second, length normalisation: a long document is expected to contain more words, so each match counts for less relative to the average document length. A short document where the term is a large share of the text looks more focused on it. Which behaviour would you prefer for your content, and why?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Build and explain an inverted index
  • Choose text analysis steps such as stemming and stop words for your content
  • Explain how TF-IDF and BM25 rank documents
  • Use field boosts, phrase queries and fuzzy matching effectively
  • Recognise vocabulary mismatch and reduce it with expansion or hybrid search

Lesson plan

6 lessons. Pick one to start there.

  1. 1 The inverted index Build a small inverted index by hand and use it to answer queries. Start
  2. 2 Text analysis Decide how text is split and normalised before indexing. Start
  3. 3 Scoring with TF-IDF and BM25 Understand how keyword engines rank matching documents. Start
  4. 4 Query features Use phrases, boosts and fuzzy matching to improve results. Start
  5. 5 Vocabulary mismatch Reduce misses caused by different words for the same thing. Start
  6. 6 Languages and evaluation Handle other languages and measure the effect of changes. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers who want to understand the search technique that still powers most search boxes and remains essential in hybrid RAG systems. You learn how an inverted index works, how text analysis (tokenising, lowercasing, stemming, stop words) shapes what matches, how TF-IDF and BM25 rank documents, and how to use field boosts, phrase queries and fuzzy matching. You also see where keyword search shines (codes, names, exact phrases) and fails (synonyms and paraphrase), and how query expansion and analysers for other languages help. Concepts apply to search engines and database full text search alike.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Emeka Nwosu

Search engineer teaching embeddings, chunking, vector and keyword search, and reranking from first principles

9 tutors 4.6(18) 374 lessons taught Sample

I come from search: indexes, ranking and the long tail of queries that make a search box look foolish. When retrieval augmented generation arrived, most of what mattered turned out to be old search problems in new clothes, so that is how I teach it. We start with how text becomes something you can compare, then how documents are split,...

See Emeka's profile and tutors