Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Evaluating Retrieval Quality

Measure whether your retrieval finds the right passages with real queries, relevance labels and the right metrics.

By Fumiko Arai Advanced Evaluation and testing RAG and search 4.7(3) 75 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $11 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Evaluating Retrieval Quality AI tutor following Fumiko Arai's plan
Student:

We switched to smaller chunks and our recall dropped, but answers seem better. How can both be true?

Tutor:

Check how relevance is labelled. If labels are tied to old chunk ids, the new smaller chunks may contain the right text but not count as relevant, so recall looks worse than it is. Relabel at the span level: a chunk is relevant if it overlaps a labelled relevant span, then recompute both setups on the same labels. Also compare with k adjusted, since smaller chunks may need a larger k for the same evidence. How were your current relevance labels created?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Build a retrieval test set from real and synthetic queries
  • Label relevance so different chunking strategies can be compared fairly
  • Choose and compute recall, precision, MRR and nDCG at the right k
  • Validate model judged relevance against human labels
  • Separate retrieval failures from generation failures end to end

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Building a query set Assemble real and synthetic queries that reflect actual use. Start
  2. 2 Labelling relevance Label relevance at a level that survives chunking changes. Start
  3. 3 Retrieval metrics Compute and interpret recall, precision, MRR and nDCG. Start
  4. 4 Model judged relevance Scale labelling with a model while keeping it honest. Start
  5. 5 Reading results Segment results and avoid drawing conclusions from noise. Start
  6. 6 Retrieval inside end to end evaluation Attribute each bad answer to retrieval or generation. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For experienced developers who tune chunking, embeddings, hybrid search or rerankers and need evidence instead of impressions. You learn to build a retrieval test set from real and synthetic queries, label relevance at the right level so different chunking strategies can be compared, and choose metrics: recall at k, precision at k, MRR and nDCG. You also check model judged relevance against human labels, segment results by query type, and separate retrieval failures from generation failures in end to end RAG evaluation. You finish with a repeatable evaluation you can run on every change.

Reviews

4.7

3 ratingsSample

  • Mateus L.Sample

    The pooling bias explanation was eye opening. We now label the union of results from both systems before comparing.

  • Cecile B.Sample

    Rigorous, as promised. Checking the model judge against our human labels showed it was too lenient, which we adjusted for.

  • Ruth A.Sample

    Our labels were tied to chunk ids, so every chunking experiment looked worse. Span level labelling fixed the comparison and changed our decision.

About the teacher

Fumiko Arai

Takes retrieval systems from demo to dependable: parsing, citations, freshness, retrieval evaluation and debugging

9 tutors 4.6(14) 267 lessons taught Sample

Most RAG demos work on the ten documents someone picked. I teach what happens after that: scanned PDFs, tables, documents that change every week, answers that cite the wrong page and users who ask things the documents never covered. My background is in document processing and internal knowledge tools, so I am practical about formats and sceptical of any setup...

See Fumiko's profile and tutors