The transformer, block by block
Trace a token through every part of a transformer and count where the parameters live
A taste of a lesson
How many parameters does one block have if the width is 1024 and the FFN expands 4x?
Split it into the two sublayers. Attention has four projection matrices, Q, K, V and output, each 1024 by 1024, so 4 times 1,048,576, about 4.2 million. The feed forward network goes 1024 to 4096 and back, two matrices of 1024 by 4096, so 8 times 1,048,576, about 8.4 million. Together that is about 12.6 million, which is the 12d squared rule, ignoring small biases and norm weights. Notice the feed forward part is twice the attention. Now you try: with 24 such blocks and a 50,000 token vocabulary with tied embeddings, what is the rough total?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Trace a token through embedding, residual stream, attention, feed forward and output head
- Estimate a transformer's parameter count from its width, depth and vocabulary
- Explain pre norm residual blocks and why the residual stream matters
- Compare encoder only, decoder only and encoder decoder designs and their uses
- Read a model configuration file and explain each architectural setting
Lesson plan
- 1 From token ids to vectors Turn token ids into embedding vectors and understand the size of the embedding table. Start
- 2 The residual stream and pre norm blocks See how each block reads from and writes back to a shared stream of vectors. Start
- 3 Attention inside the block Place multi head attention in the block and track its shapes and parameters. Start
- 4 The feed forward network Understand the per position MLP that holds most of each block's parameters. Start
- 5 Stacking blocks and the output head Go from the final hidden state to next token probabilities. Start
- 6 Encoder only, decoder only, encoder decoder Match the three transformer families to the tasks they suit. Start
Try asking
About this tutor
For learners who already understand attention and want to see the whole transformer as one working machine. You follow a single token from its embedding through the residual stream, normalisation, multi head attention and the feed forward network, then stack the blocks and reach the output head that turns vectors back into word probabilities. Along the way you track shapes, estimate parameter counts for a realistic model, and compare encoder only, decoder only and encoder decoder designs. The aim is that you can read any transformer architecture diagram or config file and say what each part does, how big it is and why it is there.
Reviews
4.7
3 ratingsSample
- Viktor H.Sample
The 12d squared rule plus the embedding table let me sanity check a config file at work in about a minute. Lesson two on the residual stream changed how I picture the whole model.
- Tariq S.Sample
Finally understand why people say the MLP layers store a lot of the knowledge. Counting parameters by hand made it obvious where the bulk is.
- Ines M.Sample
Dense but well ordered. I would have liked one more example of a gated feed forward layer, though the tutor explained the parameter trade off clearly when I asked.
About the teacher
Architectures explained from the inside: convolutions, recurrence, attention and beyond
9 tutors 362 lessons taught Sample
I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...
See Nikolai's profile and tutorsMore like this
Other tutors on the same or nearby topics.