Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

The attention formula, step by step

Work through queries, keys, values and softmax by hand until the transformer equation makes sense

By Bastian Weber Advanced How language models work 4.5(2) 48 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $9 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

The attention formula, step by step AI tutor following Bastian Weber's plan
Student:

Why do we divide QK^T by the square root of d_k? It seems arbitrary.

Tutor:

It has a clear reason. If query and key entries are roughly independent with mean 0 and variance 1, their dot product sums d_k such terms, so its variance is about d_k and its typical size grows like sqrt(d_k). With d_k = 128, raw scores might differ by tens, and softmax then puts nearly all weight on one token, with gradients close to zero elsewhere. Dividing by sqrt(d_k) brings the variance back to about 1. Exercise: for d_k = 64, what number do we divide by?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Compute single head self attention by hand on a small example
  • Explain the role of scaling, softmax and the causal mask
  • Describe multi head attention and the cost of attention in sequence length
  • Place attention inside a full transformer block with positions, residuals and normalisation

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Queries, keys and values Understand the three projections and what each represents. Start
  2. 2 Scores, scaling and softmax by hand Compute attention weights for a two token example. Start
  3. 3 Weighted values and the matrix form Finish the computation and express it compactly. Start
  4. 4 Masks and multiple heads Add causal masking and split attention into heads. Start
  5. 5 Cost, KV cache and positions Understand efficiency issues and why position must be added. Start
  6. 6 The full transformer block Assemble attention with normalisation, residuals and the feed forward layer. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For technical learners comfortable with basic algebra who want to really understand self attention. You compute attention by hand on a tiny example: build query, key and value vectors, take dot products, scale, apply softmax, and form the weighted sum. Then you scale up to matrix form, add the causal mask, split into multiple heads, and see why the cost grows with the square of sequence length. You finish by placing attention inside a full transformer block with residual connections, normalisation and the feed forward layer, and by understanding why positional information has to be added. Short pseudocode is used, no specific library.

Reviews

4.5

2 ratingsSample

  • Ilya S.Sample

    Doing the two token example by hand made the formula obvious. The variance explanation for the scaling was the best I have seen.

  • Neha G.Sample

    Rigorous and well ordered. The KV cache and grouped query section was a bit quick for me, but the shape tracing exercise helped a lot.

About the teacher

Bastian Weber

I explain how language models really work, from tokens to attention, without hand waving

9 tutors 4.5(24) 525 lessons taught Sample

I like taking the mystery out of language models. I teach what happens between typing a question and reading an answer: tokens, context windows, embeddings, attention, training and fine tuning. I start every topic with a picture or a small worked example and only add maths when it earns its place. My work background is in software and teaching workshops...

See Bastian's profile and tutors