Build Your First Eval Set
Create a small, honest set of test cases for your LLM feature and use it to judge every change.
A taste of a lesson
I changed my prompt and the five examples I tried look better. Is that enough to ship?
Five hand picked examples are a good start but easy to fool yourself with: they are probably typical cases, and outputs vary between runs. Collect around thirty cases, including tricky ones and a few your feature should decline, write what a good answer must contain for each, then run both prompt versions on all of them. Compare pass rates and read every case that changed. If the new prompt fixes three cases but breaks two, you want to know. What kinds of tricky inputs do your users send?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what an evaluation set is and why impressions mislead
- Collect 20 to 50 cases that cover typical, edge, adversarial and unanswerable inputs
- Write expected behaviour for each case in a gradeable way
- Grade with simple automatic checks and a short human rubric
- Record results by version and grow the set from real failures
Lesson plan
- 1 Why evaluate at all See how impressions mislead and what an eval set gives you. Start
- 2 Collecting cases Gather 20 to 50 cases from real or realistic use. Start
- 3 Writing expected behaviour Describe good output for each case so it can be graded. Start
- 4 Simple grading Grade with automatic checks first and human judgement where needed. Start
- 5 Running and comparing versions Record results and compare two versions on the same set. Start
- 6 Growing the set honestly Keep the set useful as the product evolves. Start
Try asking
About this tutor
For beginners who change prompts or models and judge the result by trying a few examples and getting a feeling. You learn what an evaluation set is, how to collect 20 to 50 cases from real use, what each case should contain, how to cover typical, tricky, adversarial and unanswerable inputs, and how to grade with simple automatic checks first and human judgement where needed. You run your set, record results with the prompt and model version, read the failures and grow the set from real problems. A spreadsheet is enough to start, and this tutor is free.
Reviews
4.7
3 ratingsSample
- Ana P.Sample
Had a working eval spreadsheet after the first session. It immediately showed my new prompt broke the decline behaviour I thought was fine.
- Ming W.Sample
Free and genuinely changed how I work. Reading every changed case instead of just the score is the habit I kept.
- Kevin O.Sample
Simple and encouraging. Starting with contains and JSON checks before anything fancy was the right call for a beginner.
About the teacher
Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments
9 tutors 294 lessons taught Sample
I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...
See Gonzalo's profile and tutorsMore like this
Other tutors on the same or nearby topics.