Evaluating Retrieval Quality
Measure whether your retrieval finds the right passages with real queries, relevance labels and the right metrics.
A taste of a lesson
We switched to smaller chunks and our recall dropped, but answers seem better. How can both be true?
Check how relevance is labelled. If labels are tied to old chunk ids, the new smaller chunks may contain the right text but not count as relevant, so recall looks worse than it is. Relabel at the span level: a chunk is relevant if it overlaps a labelled relevant span, then recompute both setups on the same labels. Also compare with k adjusted, since smaller chunks may need a larger k for the same evidence. How were your current relevance labels created?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Build a retrieval test set from real and synthetic queries
- Label relevance so different chunking strategies can be compared fairly
- Choose and compute recall, precision, MRR and nDCG at the right k
- Validate model judged relevance against human labels
- Separate retrieval failures from generation failures end to end
Lesson plan
- 1 Building a query set Assemble real and synthetic queries that reflect actual use. Start
- 2 Labelling relevance Label relevance at a level that survives chunking changes. Start
- 3 Retrieval metrics Compute and interpret recall, precision, MRR and nDCG. Start
- 4 Model judged relevance Scale labelling with a model while keeping it honest. Start
- 5 Reading results Segment results and avoid drawing conclusions from noise. Start
- 6 Retrieval inside end to end evaluation Attribute each bad answer to retrieval or generation. Start
Try asking
About this tutor
For experienced developers who tune chunking, embeddings, hybrid search or rerankers and need evidence instead of impressions. You learn to build a retrieval test set from real and synthetic queries, label relevance at the right level so different chunking strategies can be compared, and choose metrics: recall at k, precision at k, MRR and nDCG. You also check model judged relevance against human labels, segment results by query type, and separate retrieval failures from generation failures in end to end RAG evaluation. You finish with a repeatable evaluation you can run on every change.
Reviews
4.7
3 ratingsSample
- Mateus L.Sample
The pooling bias explanation was eye opening. We now label the union of results from both systems before comparing.
- Cecile B.Sample
Rigorous, as promised. Checking the model judge against our human labels showed it was too lenient, which we adjusted for.
- Ruth A.Sample
Our labels were tied to chunk ids, so every chunking experiment looked worse. Span level labelling fixed the comparison and changed our decision.
About the teacher
Takes retrieval systems from demo to dependable: parsing, citations, freshness, retrieval evaluation and debugging
9 tutors 267 lessons taught Sample
Most RAG demos work on the ten documents someone picked. I teach what happens after that: scanned PDFs, tables, documents that change every week, answers that cite the wrong page and users who ask things the documents never covered. My background is in document processing and internal knowledge tools, so I am practical about formats and sceptical of any setup...
See Fumiko's profile and tutorsMore like this
Other tutors on the same or nearby topics.