Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Debugging a RAG App That Answers Badly

Trace each bad RAG answer to its real cause, from missing documents to ignored passages, and fix the most common.

By Fumiko Arai Intermediate RAG and search 4.7(3) 56 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $8 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Debugging a RAG App That Answers Badly AI tutor following Fumiko Arai's plan
Student:

Our bot answers the first question well, but 'and what about part-time staff?' gets nonsense. Why?

Tutor:

Look at what was actually searched for. A follow up like that, searched on its own, has almost no meaning: no mention of the topic from the first question, such as holiday entitlement. Retrieval then returns random part time content. The usual fix is to rewrite the follow up into a standalone query using the conversation history before searching, for example 'holiday entitlement for part time staff'. Then log the rewritten query to check it. What query does your trace show for that turn?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Reproduce a bad answer from a complete trace
  • Walk a fixed cause list from content gaps to citation errors
  • Tell retrieval failures from generation failures with quick checks
  • Keep a categorised failure log and fix the largest category first
  • Turn fixed failures into permanent evaluation cases

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Reproduce from a full trace Capture everything needed to replay a bad answer. Start
  2. 2 Content and parsing problems Rule out missing documents and broken extraction first. Start
  3. 3 Retrieval misses Find out why the right chunk was not retrieved. Start
  4. 4 Ranking and generation problems Diagnose good evidence that was cut off, ignored or misread. Start
  5. 5 Failure logs and priorities Categorise failures and fix the most common cause first. Start
  6. 6 Keeping fixes fixed Protect improvements with evaluation cases. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers whose retrieval augmented app gives wrong, vague or unsupported answers and who are tired of tweaking prompts at random. You learn a systematic method: reproduce the failure from a full trace, then walk down a fixed list of causes, from the answer not existing in the corpus, through parsing, chunking, filters and ranking, to the model ignoring or misreading good evidence. You practise on real failure cases, keep a failure log with categories and counts, fix the most frequent cause first, and turn every fixed case into a permanent evaluation example so it stays fixed.

Reviews

4.7

3 ratingsSample

  • Darius N.Sample

    The trick of rerunning with only the correct passage told us immediately whether to fix retrieval or generation. Practical lessons.

  • Lucia M.Sample

    The cause list in order is now pinned in our team channel. Half our bad answers were content gaps, not model problems. We stopped rewriting the prompt every day.

  • Sophie R.Sample

    Follow up question rewriting fixed our worst complaints. The failure log habit keeps us honest about priorities.

About the teacher

Fumiko Arai

Takes retrieval systems from demo to dependable: parsing, citations, freshness, retrieval evaluation and debugging

9 tutors 4.6(14) 267 lessons taught Sample

Most RAG demos work on the ten documents someone picked. I teach what happens after that: scanned PDFs, tables, documents that change every week, answers that cite the wrong page and users who ask things the documents never covered. My background is in document processing and internal knowledge tools, so I am practical about formats and sceptical of any setup...

See Fumiko's profile and tutors