RNNs, LSTMs and why transformers took over
Understand recurrent networks, their gates and limits, and the real reasons attention replaced them
A taste of a lesson
If LSTMs fixed vanishing gradients, why did people still move to transformers?
LSTMs eased the problem rather than removing it, and two other limits remained. First, information between distant tokens still has to pass through many sequential steps, so long range links stay hard to learn. Second, and practically decisive, an LSTM processes tokens one after another, so training cannot run in parallel across the sequence. Transformers connect every pair of positions in a single attention step and train all positions at once, which fits GPUs and large datasets far better. The cost is attention that grows with the square of sequence length. Quick check: at inference, which model needs memory that grows with the length of the text so far?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Trace a simple recurrent network over a short sequence by hand
- Explain vanishing gradients through time and how LSTM and GRU gates help
- Describe how attention fixed the bottleneck in encoder decoder models
- Weigh the real trade offs between recurrent models and transformers
- Recognise where recurrence and newer recurrent style models still fit
Lesson plan
- 1 Recurrence by hand Step a tiny RNN through a three token sequence and see how state carries information. Start
- 2 Backpropagation through time Understand why gradients vanish or explode across long sequences. Start
- 3 LSTM and GRU gates Learn how gates control what to keep, forget and output. Start
- 4 Encoder decoder models and attention See how attention removed the fixed vector bottleneck in translation. Start
- 5 Why transformers took over Compare recurrent models and transformers on training speed, memory and range. Start
- 6 Where recurrence still lives Identify current uses of recurrence and the research reviving it. Start
Try asking
About this tutor
For learners who know basic neural networks and want the story of sequence models told properly. You will step a simple recurrent network through a short sequence by hand, see why its gradients fade over long spans, and learn how LSTM and GRU gates work around that. Then you follow the encoder decoder models used for translation, the attention mechanism added to them, and the shift to transformers, weighing the real trade offs: parallel training and short paths between positions against quadratic attention cost. The course ends with where recurrence still makes sense and with newer recurrent style architectures presented as active research, not settled fact.
Reviews
4.5
2 ratingsSample
- Laura B.Sample
The hand traced RNN in lesson one made backprop through time make sense. Before this, LSTM diagrams were just boxes and arrows to me.
- Samir H.Sample
Balanced and clear. I appreciated that state space models were presented as research rather than the next big thing.
About the teacher
Architectures explained from the inside: convolutions, recurrence, attention and beyond
9 tutors 362 lessons taught Sample
I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...
See Nikolai's profile and tutorsMore like this
Other tutors on the same or nearby topics.