Optimisation Basics for ML
Understand momentum, adaptive methods, schedules and the shape of loss landscapes
A taste of a lesson
My loss drops nicely for 200 steps, then spikes and goes to NaN. Learning rate is constant. What should I check?
A late spike to NaN usually means an update became too large somewhere. Common causes: the learning rate is near the edge of stability and a steep region or unusual batch pushes it over; gradients exploded through depth; or a numerical overflow, such as an exponential in a loss computed without a stable formulation. Try, one at a time: log the gradient norm to see if it jumps before the spike; add gradient clipping; lower the learning rate or add decay. Does the gradient norm rise gradually or suddenly before the NaN?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain convexity, conditioning and saddle points and their effects
- Derive momentum and adaptive updates from the problems they solve
- Design learning rate warmup and decay schedules
- Apply weight decay correctly with adaptive optimisers
- Diagnose training curves and choose fixes
Lesson plan
- 1 Loss landscapes Understand convexity, conditioning and saddle points. Start
- 2 Momentum Derive momentum as a running velocity that smooths updates. Start
- 3 Adaptive methods Understand per parameter step scaling and its trade offs. Start
- 4 Schedules Design warmup and decay schedules for stable training. Start
- 5 Weight decay and clipping Regularise and stabilise updates correctly. Start
- 6 Reading training curves Diagnose problems from loss behaviour. Start
Try asking
About this tutor
An advanced tutor for learners who know gradient descent and want to understand the optimisers used to train real models. You will study convexity and conditioning, saddle points in high dimensions, momentum, adaptive per parameter methods, learning rate warmup and decay, weight decay and its interaction with adaptive optimisers, gradient clipping and the role of batch size. Each method is derived from the problem it solves and examined for its failure modes. The tutor stays independent of any framework and teaches the reasoning needed to debug training curves.
Reviews
4.7
3 ratingsSample
- Mariana C.Sample
Training curve diagnostics were excellent. Dense material, so I took it slowly over a few weeks.
- Jakob W.Sample
The 'each optimiser fixes a weakness' framing made everything coherent. Decoupled weight decay finally makes sense.
- Wei Z.Sample
Rigorous and framework neutral. Helped me debug a divergence problem in my own model.
About the teacher
Calculus and optimisation for machine learning, done by hand
9 tutors 319 lessons taught Sample
I teach the calculus and optimisation that make models learn: derivatives, gradients, the chain rule, gradient descent and the loss functions it minimises. My background is in engineering and numerical computing, so I care about why things work and also about when they break, such as unstable learning rates or overflowing exponentials. I teach with pencil calculations first, then a...
See Leandro's profile and tutorsMore like this
Other tutors on the same or nearby topics.