Optimisers: SGD, momentum and Adam
Choose and tune optimisers and learning rate schedules with understanding instead of guesswork
A taste of a lesson
My transformer's loss drops for 200 steps, then suddenly jumps to NaN. I use Adam with lr 1e-3 and no schedule.
That pattern often points to a learning rate that is too high for this model, especially without warmup. Early in training, Adam's step size estimates are unreliable and transformers are sensitive to large updates, so a spike can push activations into overflow. Try three changes, one at a time: add a linear warmup over the first few hundred steps, lower the peak learning rate by a factor of 3 to 10, and clip gradients by global norm. Log the gradient norm too; a sharp rise just before the NaN supports this diagnosis. What does your gradient norm do around step 200?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute update steps for SGD, momentum and Adam on a small example
- Explain why the learning rate is usually the most important hyperparameter
- Configure warmup, decay schedules and weight decay sensibly
- Read a loss curve and form a hypothesis about optimiser settings
- Design a fair comparison between optimiser settings
Lesson plan
- 1 Gradient descent and the learning rate See how step size controls whether training converges, crawls or diverges. Start
- 2 Momentum Understand how a running velocity smooths noisy gradients and speeds up progress. Start
- 3 Adaptive methods: RMSProp and Adam Learn how per parameter step sizes work and what Adam's settings mean. Start
- 4 Weight decay and AdamW Explain why decoupled weight decay differs from an L2 penalty under Adam. Start
- 5 Schedules, warmup and batch size Set a learning rate schedule and adjust it when batch size changes. Start
- 6 Tuning and comparing fairly Run a learning rate range test and compare settings without fooling yourself. Start
Try asking
About this tutor
For learners who train models and want to understand what the optimiser setting actually does. Starting from plain stochastic gradient descent, you build up momentum, per parameter step sizes and the Adam family, computing one or two update steps by hand for each. Then you learn the practical side: why the learning rate matters more than almost anything else, how warmup and decay schedules help, what weight decay really means in AdamW, how batch size interacts with learning rate, and how to run a fair comparison between settings. The aim is that you can look at a loss curve and form a sensible hypothesis about the optimiser.
Reviews
4.7
3 ratingsSample
- Aisha B.Sample
The AdamW lesson finally explained why weight decay and L2 are not the same thing with Adam. The hand computed momentum steps were a bit tedious but helped.
- Chen W.Sample
I loved the lesson on fair comparisons. My team had been declaring winners from single runs.
- Marek L.Sample
Warmup plus a lower peak learning rate fixed my exploding fine tune. More importantly I understand why now, instead of copying configs.
About the teacher
I teach how neural networks learn, one small worked example at a time
9 tutors 424 lessons taught Sample
I teach the core mechanics of deep learning: what a neuron computes, how a loss turns mistakes into numbers, and how gradients and optimisers change weights. My background is in building and training models for applied research teams, which mostly meant staring at loss curves that refused to go down. That shaped how I teach. I start every idea with...
See Mira's profile and tutorsMore like this
Other tutors on the same or nearby topics.