Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

RNNs, LSTMs and why transformers took over

Understand recurrent networks, their gates and limits, and the real reasons attention replaced them

By Nikolai Sorin Intermediate Deep learning NLP 4.5(2) 32 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

RNNs, LSTMs and why transformers took over AI tutor following Nikolai Sorin's plan
Student:

If LSTMs fixed vanishing gradients, why did people still move to transformers?

Tutor:

LSTMs eased the problem rather than removing it, and two other limits remained. First, information between distant tokens still has to pass through many sequential steps, so long range links stay hard to learn. Second, and practically decisive, an LSTM processes tokens one after another, so training cannot run in parallel across the sequence. Transformers connect every pair of positions in a single attention step and train all positions at once, which fits GPUs and large datasets far better. The cost is attention that grows with the square of sequence length. Quick check: at inference, which model needs memory that grows with the length of the text so far?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Trace a simple recurrent network over a short sequence by hand
  • Explain vanishing gradients through time and how LSTM and GRU gates help
  • Describe how attention fixed the bottleneck in encoder decoder models
  • Weigh the real trade offs between recurrent models and transformers
  • Recognise where recurrence and newer recurrent style models still fit

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Recurrence by hand Step a tiny RNN through a three token sequence and see how state carries information. Start
  2. 2 Backpropagation through time Understand why gradients vanish or explode across long sequences. Start
  3. 3 LSTM and GRU gates Learn how gates control what to keep, forget and output. Start
  4. 4 Encoder decoder models and attention See how attention removed the fixed vector bottleneck in translation. Start
  5. 5 Why transformers took over Compare recurrent models and transformers on training speed, memory and range. Start
  6. 6 Where recurrence still lives Identify current uses of recurrence and the research reviving it. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For learners who know basic neural networks and want the story of sequence models told properly. You will step a simple recurrent network through a short sequence by hand, see why its gradients fade over long spans, and learn how LSTM and GRU gates work around that. Then you follow the encoder decoder models used for translation, the attention mechanism added to them, and the shift to transformers, weighing the real trade offs: parallel training and short paths between positions against quadratic attention cost. The course ends with where recurrence still makes sense and with newer recurrent style architectures presented as active research, not settled fact.

Reviews

4.5

2 ratingsSample

  • Laura B.Sample

    The hand traced RNN in lesson one made backprop through time make sense. Before this, LSTM diagrams were just boxes and arrows to me.

  • Samir H.Sample

    Balanced and clear. I appreciated that state space models were presented as research rather than the next big thing.

About the teacher

Nikolai Sorin

Architectures explained from the inside: convolutions, recurrence, attention and beyond

9 tutors 4.6(22) 362 lessons taught Sample

I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...

See Nikolai's profile and tutors