Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Evaluating generated text

Measure the quality of generated text with metrics, people and model judges, and know each one's limits

Your first 3 replies from the tutor are free. Then you can continue the lesson for $8 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Evaluating generated text AI tutor following Mateo Rojas's plan
Student:

I used a model judge to compare two prompts. Prompt B won 58 out of 100. Is B better?

Tutor:

Not yet clear. Two things to check first. Noise: with 100 comparisons, a 58 percent win rate has a rough 95 percent interval of about 48 to 68 percent, which includes 50, so it could be chance. Bias: did you swap the order of A and B for each pair, and are B's answers longer? Judges often favour the second position or longer text. Rerun with swapped order, compare answer lengths, and have a person rate 30 pairs to see if they agree with the judge. Quick check: what would you conclude if B wins only when it appears second?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain what BLEU, ROUGE, chrF and embedding metrics measure and where they fail
  • Design a human evaluation with a clear rubric, pairwise comparisons and agreement checks
  • Use a model as judge with bias mitigations and calibration against human ratings
  • Decide whether a difference between systems is larger than noise
  • Keep test sets fresh and free from training contamination

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Start from the decision Choose an evaluation approach based on what you need to decide and what can go wrong. Start
  2. 2 Reference based metrics Understand BLEU, ROUGE, chrF and embedding similarity and their blind spots. Start
  3. 3 Human evaluation that holds up Design rubrics, comparisons and agreement checks that produce reliable judgments. Start
  4. 4 Model as judge Use model graders at scale while controlling their known biases. Start
  5. 5 Task specific checks Build targeted automatic checks that test what actually matters for the task. Start
  6. 6 Noise, sample size and contamination Report results with uncertainty and keep test data trustworthy. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who generate text with models and need to know whether a change made things better. You will learn what reference based metrics such as BLEU, ROUGE, chrF and embedding similarity actually measure and where they mislead, then design human evaluations with clear rubrics, pairwise comparisons and agreement checks. A full lesson covers using a model as a judge, including position bias, length bias and self preference, and how to calibrate a judge against human ratings. You finish with task specific checks, sample sizes and confidence intervals, and the problem of test data leaking into training. The emphasis is on choosing evaluation that matches the decision you need to make.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Mateo Rojas

Generative models and representation learning, explained with intuition first and maths second

9 tutors 4.6(20) 335 lessons taught Sample

I teach how models learn useful representations and how they generate new data: autoencoders, GANs, diffusion models, self supervised learning and language model pretraining. I came to this through research engineering work where we had to decide which kind of model was worth the compute, so I teach with trade offs in mind. Each topic starts with a picture or...

See Mateo's profile and tutors