Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.

Teacher since June 2026

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9

tutors built

4.5Sample

average from 17 reviews

294Sample

lessons taught by their tutors

About Gonzalo

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in a way you can defend, and keep checking as prompts and models change. I am wary of single scores and of graders nobody has checked, and I try to make evaluation feel like a habit rather than a big project.

Knows about

  • eval set design
  • golden answers
  • rubric grading
  • model graded evaluation
  • prompt regression tests
  • hallucination measurement
  • human review
  • A/B testing
  • red teaming and prompt injection testing

Tutors by Gonzalo

9 tutors

Model Graded Evals and Their Pitfalls

Model Graded Evals and Their Pitfalls

Use language models as graders without fooling yourself: biases, validation against people and safeguards.AdvancedEvaluation and testing4.3(4)60 lessonsSample
Gonzalo Ibarra$11
Build Your First Eval Set

Build Your First Eval Set

Create a small, honest set of test cases for your LLM feature and use it to judge every change.BeginnerEvaluation and testing4.7(3)52 lessonsSample
Gonzalo IbarraFree
Red Teaming and Prompt Injection Testing

Red Teaming and Prompt Injection Testing

Test your LLM app against jailbreaks, prompt injection and data leaks with a repeatable attack suite.AdvancedEvaluation and testing4.7(3)49 lessonsSample
Gonzalo Ibarra$12
Rubric Grading for Open Ended Output

Rubric Grading for Open Ended Output

Design rubrics that make grading emails, summaries and explanations consistent, fair and repeatable.IntermediateEvaluation and testing4.3(3)49 lessonsSample
Gonzalo Ibarra$6
Measuring Hallucination in LLM Apps

Measuring Hallucination in LLM Apps

Measure how often your app states unsupported things, with claim level checks you can repeat and trust.IntermediateEvaluation and testing4.5(2)42 lessonsSample
Gonzalo Ibarra$8
Regression Testing Your Prompts

Regression Testing Your Prompts

Catch quality drops before users do by running an eval suite on every prompt, model or setting change.IntermediateEvaluation and testing4.5(2)42 lessonsSample
Gonzalo Ibarra$7
Golden Answers and Reference Checks

Golden Answers and Reference Checks

Write reference answers experts agree on and compare model output to them with the right matching method.BeginnerEvaluation and testingNew
Gonzalo Ibarra$5
Human Review That Scales

Human Review That Scales

Set up human review of AI outputs that is consistent, affordable and feeds back into better systems.All levelsEvaluation and testingNew
Gonzalo Ibarra$5
A/B Testing LLM Features in Production

A/B Testing LLM Features in Production

Run fair online experiments on prompts, models and LLM features and read the results without fooling yourself.AdvancedEvaluation and testingNew
Gonzalo Ibarra$10

Recent reviews

What students said about Gonzalo's tutors.

  • Ola N.Sample

    The faithfulness versus factuality distinction cleared up weeks of confused discussion. Validating the model judge took effort but was necessary.

    On Measuring Hallucination in LLM Apps

  • Hassan B.Sample

    Diff reports with outputs side by side changed our review meetings. We caught a prompt edit that fixed tone but broke refusals in one category.

    On Regression Testing Your Prompts

  • Beatriz A.Sample

    Good for domain experts like me who are not engineers. Kappa was explained simply enough to use.

    On Rubric Grading for Open Ended Output

  • Mia K.Sample

    Clear that prompt patches are not real fixes. We reduced tool permissions instead and the attack success rate dropped across categories.

    On Red Teaming and Prompt Injection Testing

  • Imani W.Sample

    Claim level labelling gave us a number we could defend to leadership, with sample size and all. The unanswerable question set was the most revealing part.

    On Measuring Hallucination in LLM Apps

  • Ana P.Sample

    Had a working eval spreadsheet after the first session. It immediately showed my new prompt broke the decline behaviour I thought was fine.

    On Build Your First Eval Set