Train, Validation and Test Splits
Split your data so your model's score means something outside your laptop
A taste of a lesson
I scaled my whole dataset first, then split it. My teacher says that is wrong. Why does it matter?
When you scale the whole dataset, the mean and standard deviation include the test rows. So the training data was transformed using information from the test set, a small leak of the future into the past. With scaling the effect is often small, but the same habit with imputation, target encoding or feature selection can inflate scores badly. The fix is to split first, fit the scaler on training data only, then apply it to validation and test. Can you name one other preprocessing step that should follow the same rule?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain the separate jobs of training, validation and test sets
- Choose split sizes that suit the amount of data
- Use stratified, time based and grouped splits where they fit
- Fit preprocessing on training data only
- Diagnose a split that gives misleadingly good results
Lesson plan
- 1 Three sets, three jobs Understand what training, validation and test data are each used for. Start
- 2 How much data in each Pick split proportions sensibly for small and large datasets. Start
- 3 Stratified and grouped splits Keep class balance and keep related rows together. Start
- 4 Splitting by time Evaluate forecasting style models on future periods. Start
- 5 Preprocessing without peeking Fit scaling, imputation and encoding only on training data. Start
- 6 Diagnosing a bad split Recognise signs that a split is making results look too good. Start
Try asking
About this tutor
A free beginner tutor on the single most important habit in machine learning: keeping data for honest evaluation. You will learn what training, validation and test sets are for, how big each should be, and why the test set must be touched as rarely as possible. Lessons cover stratified splits, time based splits, grouped splits when the same customer or patient appears many times, and fitting preprocessing on training data only. Each lesson includes a short scenario where a split went wrong and you diagnose why. Useful before any other modelling course.
Reviews
4.7
3 ratingsSample
- Kemal Y.Sample
The grouped split lesson explained why my model looked amazing in testing and poor in production: the same customers were in both sets.
- Hannah W.Sample
Free, short and exactly what our analysts needed. The 'sealed exam paper' rule is now on our team wiki.
- Rohan P.Sample
Clear explanations. The time series lesson was brief but it pointed me in the right direction.
About the teacher
Model evaluation you can trust: splits, metrics, leakage and tuning
9 tutors 439 lessons taught Sample
Most of the machine learning failures I have seen were not about the algorithm. They came from a test set that was used too often, a feature that leaked the answer, or a metric that hid the real cost of mistakes. I teach the unglamorous discipline that makes model results believable: how to split data, how to validate, which metric...
See Lukas's profile and tutorsMore like this
Other tutors on the same or nearby topics.