Overfitting and Regularisation
Diagnose overfitting with learning curves and fix it with the right kind of regularisation
A taste of a lesson
Training R squared 0.98, validation 0.61. I added lasso and validation went to 0.64. Is that the best I can do?
Maybe not. A gap that large suggests high variance, and lasso helped a little, but first check three things. Did you standardise features before the penalty? Did you tune the penalty strength over a wide log range with cross validation? And plot a learning curve: if validation keeps rising as you add rows, more data is your best lever. Also look for leakage or entity overlap that might make training look unrealistically good. Which of these can you check first with what you already have?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Recognise overfitting and underfitting from training and validation scores
- Read learning curves and validation curves to choose a fix
- Apply ridge, lasso and elastic net with properly scaled features
- Choose regularisation strength with cross validation
- Decide when more data or better features beat regularisation
Lesson plan
- 1 Seeing overfitting Spot overfitting and underfitting by comparing training and validation error. Start
- 2 Bias and variance in pictures Connect bias and variance to model flexibility with a polynomial example. Start
- 3 Learning and validation curves Use curves to decide between more data, simpler models or better features. Start
- 4 Ridge, lasso and elastic net Apply penalty based regularisation to linear models correctly. Start
- 5 Regularisation beyond linear models Recognise the same idea in trees, boosting and neural networks. Start
- 6 When regularisation is not the fix Separate overfitting from leakage, bad splits and weak features. Start
Try asking
About this tutor
An intermediate tutor on the tension at the heart of machine learning: fitting the data you have without memorising it. You will read learning curves and validation curves, understand bias and variance without hand waving, and apply ridge, lasso and elastic net penalties with properly scaled features. Lessons extend the idea to trees, boosting and early stopping, and discuss when more data helps more than any penalty. Exercises use small polynomial fits and realistic tabular models so you can see overfitting appear and disappear as you change one setting at a time.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Model evaluation you can trust: splits, metrics, leakage and tuning
9 tutors 439 lessons taught Sample
Most of the machine learning failures I have seen were not about the algorithm. They came from a test set that was used too often, a feature that leaked the answer, or a metric that hid the real cost of mistakes. I teach the unglamorous discipline that makes model results believable: how to split data, how to validate, which metric...
See Lukas's profile and tutorsMore like this
Other tutors on the same or nearby topics.