Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Rubric Grading for Open Ended Output

Design rubrics that make grading emails, summaries and explanations consistent, fair and repeatable.

By Gonzalo Ibarra Intermediate Evaluation and testing 4.3(3) 49 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $6 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Rubric Grading for Open Ended Output AI tutor following Gonzalo Ibarra's plan
Student:

Our rubric has 'tone is appropriate' and reviewers keep disagreeing on it. How do I fix that?

Tutor:

Appropriate means different things to different people, so make it observable. Ask your reviewers what they actually look for and write those as separate checks, for example: greets the customer by name; no blame language such as you should have; no internal jargon; apologises once at most. Add a pass and a fail example for each. Then have two reviewers grade the same fifteen replies and compare. Which disagreements came up most often before?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Turn vague quality goals into specific, observable criteria
  • Choose between pass or fail checks, anchored scales and pairwise comparisons
  • Set must pass criteria for accuracy, safety and policy
  • Calibrate graders and measure agreement between them
  • Report results per criterion and prepare rubrics for model graders

Lesson plan

6 lessons. Pick one to start there.

  1. 1 From purpose to criteria Derive criteria from what makes outputs useful or harmful. Start
  2. 2 Choosing scales Pick pass or fail, anchored scales or pairwise judgement per criterion. Start
  3. 3 Must pass criteria Make accuracy, safety and policy failures impossible to average away. Start
  4. 4 Grader instructions Write instructions that let different people grade the same way. Start
  5. 5 Calibration and agreement Align graders and measure how consistently they judge. Start
  6. 6 Reporting and reuse Report results usefully and prepare for automated grading. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers, product people and domain experts who need to judge outputs that have no single right answer: drafted emails, summaries, explanations, support replies. You learn to turn vague goals like good tone into specific, observable criteria, choose between pass or fail checks, anchored scales and pairwise comparisons, set must pass criteria for safety and accuracy, write grader instructions with examples, and calibrate graders until they agree. You measure agreement between graders, report results per criterion and prepare rubrics that can later be used by a model grader after validation.

Reviews

4.3

3 ratingsSample

  • Beatriz A.Sample

    Good for domain experts like me who are not engineers. Kappa was explained simply enough to use.

  • Rashid M.Sample

    Clear on why 1 to 10 scales fail. Pairwise comparison made our prompt version decisions much easier.

  • Johanna F.Sample

    Splitting appropriate tone into four observable checks took our reviewer agreement from frustrating to fine. The calibration exercise was worth the time.

About the teacher

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9 tutors 4.5(17) 294 lessons taught Sample

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...

See Gonzalo's profile and tutors