Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Attention mechanisms, step by step

Compute attention by hand, then understand masks, heads, KV caching and efficient variants

By Nikolai Sorin Intermediate Deep learning How language models work 4.3(4) 66 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $8 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Attention mechanisms, step by step AI tutor following Nikolai Sorin's plan
Student:

Why do we need a causal mask? Could the model not just learn not to look at future tokens?

Tutor:

During training, the whole sentence is fed in at once so every position can be trained in parallel. Without a mask, the token at position 3 could simply read position 4, which is exactly the word it is supposed to predict, so training would be trivially easy and useless. The model cannot learn to avoid this, because peeking always lowers the loss. The mask enforces the rule by setting scores for future positions to negative infinity, so softmax gives them weight zero. At generation time the future does not exist yet, so training and use then match. Quick check: which entries of a 4 by 4 score matrix does a causal mask block?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Compute scaled dot product attention by hand for a few tokens
  • Explain the roles of queries, keys, values and the square root scaling
  • Apply causal masks and describe self attention versus cross attention
  • Explain multi head attention and the key value cache used in generation
  • Compare the memory and speed trade offs of efficient attention variants

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Attention as a weighted average Compute an attention output from scores, softmax weights and values. Start
  2. 2 Queries, keys and values Derive Q, K and V from token vectors with learned projections and compute scores. Start
  3. 3 Masks, self attention and cross attention Use masks for generation and distinguish where queries, keys and values come from. Start
  4. 4 Multiple heads Understand why attention is split into heads and how their outputs are combined. Start
  5. 5 Generation and the KV cache See how cached keys and values make token by token generation affordable. Start
  6. 6 Making attention cheaper Compare grouped query attention, fast exact kernels and approximate attention. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For learners who have heard that attention is the heart of modern AI and want to actually compute it. You start with two or three tokens and work through queries, keys, values, scaled scores and softmax weights by hand, so the formula becomes a sequence of simple steps. Then you add self attention versus cross attention, causal masks for text generation, multiple heads, and the key value cache that makes generation practical. The final lessons cover efficiency: why attention cost grows with the square of sequence length, how shared key value heads reduce memory, and what fast exact kernels change. Comfort with vectors and dot products is enough.

Reviews

4.3

4 ratingsSample

  • Priya S.Sample

    The KV cache lesson explained why my local model runs out of memory with long chats. Estimating the cache size by hand was eye opening.

  • Jonas W.Sample

    Clear and precise. The efficient attention lesson was dense; I needed to ask for the difference between exact and approximate methods twice, but got there.

  • Daniel E.Sample

    Computing attention with two dimensional vectors was the key. I had read the formula many times but never actually done it.

  • Mei C.Sample

    Good material, though I already knew the basics so the first two lessons felt slow. The later lessons were worth it.

About the teacher

Nikolai Sorin

Architectures explained from the inside: convolutions, recurrence, attention and beyond

9 tutors 4.6(22) 362 lessons taught Sample

I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...

See Nikolai's profile and tutors