Evaluating a model after fine tuning
Prove a tuned model is actually better, on your task and everywhere else it matters
A taste of a lesson
My tuned model scores 84 percent versus 81 percent for the prompted base model on 100 test questions. Can I ship it?
Not on that evidence alone. With 100 questions, each score carries roughly plus or minus 8 points of uncertainty, so a 3 point gap could easily be chance. Two things would help. First, a paired comparison: on how many individual questions does the tuned model win, lose or tie against the base? That is more sensitive than comparing totals. Second, more test items, ideally drawn from real traffic. Also run your regression suite, since a small gain is not worth a safety or formatting loss. Quick exercise: if the models disagree on 20 questions, what split would start to convince you?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Build a held out task set and baselines that a tuned model must beat
- Add regression checks for general ability, safety and formatting
- Validate model graded evaluation against a sample of human judgements
- Estimate uncertainty and avoid over reading small differences
- Write a decision record that someone else could audit
Lesson plan
- 1 Defining success before you train Write down the task metric and build a held out set from real inputs. Start
- 2 Baselines that keep you honest Compare the tuned model with prompting, retrieval and the current production model. Start
- 3 Regression suites Check that tuning did not damage general ability, safety or formatting. Start
- 4 Model graded and human evaluation Use automatic graders responsibly and back them with blind human review. Start
- 5 Uncertainty and contamination Quantify how much a score could move by chance and find leaked test items. Start
- 6 Choosing a checkpoint and recording the decision Select a model on validation data and document the evidence for the choice. Start
Try asking
About this tutor
For anyone who has fine tuned a model, or is about to, and needs an honest answer to the question 'did it work?'. You will build a held out task set from real inputs, set up baselines that a tuned model must beat, and add regression suites for general ability, safety and formatting. The lessons cover model graded evaluation and how to check it against human judgement, blind human review, measuring uncertainty on small test sets, catching contamination between training and test data, and choosing checkpoints without peeking at the test set. You finish by writing a short decision record that someone else could audit.
Reviews
4.7
3 ratingsSample
- Julia P.Sample
Writing the decision record felt like homework, but my manager read it and approved the launch without a single extra meeting.
- Omar S.Sample
Strong on contamination. We found eleven near duplicates between our training and test sets. A bit more on rubric writing would have been welcome.
- Hannah G.Sample
The paired comparison idea changed our review meetings. We stopped arguing about two point differences and started looking at which questions flipped.
About the teacher
Fine tuning with judgment: when to do it, how to do it well, and how to know it worked
9 tutors 428 lessons taught Sample
I teach fine tuning and post training: choosing between prompting, retrieval and tuning, building datasets, parameter efficient methods, instruction and preference tuning, and evaluating the result. My background is in applied machine learning projects where the expensive mistake was usually tuning a model before anyone had defined what better meant. That is why I start every topic with the evaluation...
See Neha's profile and tutorsMore like this
Other tutors on the same or nearby topics.