Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

A/B Testing LLM Features in Production

Run fair online experiments on prompts, models and LLM features and read the results without fooling yourself.

Your first 3 replies from the tutor are free. Then you can continue the lesson for $10 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

A/B Testing LLM Features in Production AI tutor following Gonzalo Ibarra's plan
Student:

After three days our new model shows a significant lift in thumbs up. Can we stop and ship?

Tutor:

I would not yet. Stopping as soon as a result looks significant inflates false positives, three days misses weekly patterns, and novelty can boost early reactions. Thumbs are also a thin signal: few users click them, and those who do are not typical. Check your pre registered primary metric, such as resolution without escalation, and guardrails like latency and cost, then run to the planned end date. What primary metric did you set before launch?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Decide when an online experiment is needed after offline evals
  • Randomise by user and write a plan with primary and guardrail metrics
  • Interpret thumbs, edits, regenerations and escalations carefully
  • Plan sample size and avoid peeking, novelty and multiple comparison traps
  • Use shadow testing, gradual ramps and interleaving where they fit

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Offline versus online Understand what online experiments add to offline evaluation. Start
  2. 2 Designing the experiment Write a pre registered plan with randomisation and decision rules. Start
  3. 3 Metrics for LLM features Pick and define signals that reflect real value. Start
  4. 4 Sample size and duration Plan how many users and how long the experiment needs. Start
  5. 5 Reading results honestly Avoid the common traps when analysing outcomes. Start
  6. 6 Rolling out safely Ramp changes and choose alternatives where they fit better. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For experienced engineers, data scientists and product managers who have passed offline evaluation and now need to know whether a change helps real users. You learn when an online experiment is needed, how to randomise at the user level, choose a primary metric tied to user value and guardrail metrics for latency, cost and safety, interpret LLM specific signals such as thumbs, edits, regenerations and escalations, plan sample size, avoid peeking and novelty effects, and ramp changes safely. You also learn alternatives for ranking features, such as interleaving, and when shadow testing is the better first step.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9 tutors 4.5(17) 294 lessons taught Sample

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...

See Gonzalo's profile and tutors