Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Measuring Hallucination in LLM Apps

Measure how often your app states unsupported things, with claim level checks you can repeat and trust.

By Gonzalo Ibarra Intermediate Evaluation and testing 4.5(2) 42 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $8 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Measuring Hallucination in LLM Apps AI tutor following Gonzalo Ibarra's plan
Student:

The answer included a true fact that is not in our documents. Does that count as a hallucination?

Tutor:

For a document grounded app, count it as unfaithful, even though it is true. Your users cannot verify it from the cited sources, and next time the model adds something from general knowledge it may be outdated or wrong for your organisation. Label it not in source rather than contradicted, so you can track both kinds separately. Then decide whether outside knowledge is ever allowed, and say so in the prompt. Is outside knowledge ever acceptable in your app?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Distinguish faithfulness from factuality and intrinsic from extrinsic errors
  • Split answers into claims and label support against sources
  • Measure honest abstention on unanswerable questions
  • Validate model judges and report rates with sample sizes
  • Test hallucination reduction levers by re-measuring

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Defining what you measure Separate faithfulness, factuality and the kinds of unsupported claims. Start
  2. 2 Claims as the unit Split answers into atomic claims with consistent guidelines. Start
  3. 3 Labelling support Judge each claim against the sources reliably. Start
  4. 4 Unanswerable questions Measure whether the app admits when it does not know. Start
  5. 5 Caveats and reporting Report results honestly with their uncertainty. Start
  6. 6 Testing improvements Try reduction levers and confirm their effect. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers and evaluators who need a real number, not an anecdote, for how often their assistant makes things up. You learn to separate faithfulness (staying true to the provided sources) from factuality (being true about the world), split answers into individual claims, check each claim against sources with people or validated model judges, and measure behaviour on unanswerable questions. You also learn the caveats: hallucination rates depend heavily on task and question mix, judges make mistakes, and general published rates say little about your app. You finish with a repeatable measurement and a list of levers to test.

Reviews

4.5

2 ratingsSample

  • Ola N.Sample

    The faithfulness versus factuality distinction cleared up weeks of confused discussion. Validating the model judge took effort but was necessary.

  • Imani W.Sample

    Claim level labelling gave us a number we could defend to leadership, with sample size and all. The unanswerable question set was the most revealing part.

About the teacher

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9 tutors 4.5(17) 294 lessons taught Sample

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...

See Gonzalo's profile and tutors