Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.

Teacher since January 2026

Nikolai Sorin

Architectures explained from the inside: convolutions, recurrence, attention and beyond

9

tutors built

4.6Sample

average from 22 reviews

362Sample

lessons taught by their tutors

About Nikolai

I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and graph networks feel like a sequence of sensible decisions rather than a zoo of names. I draw a lot of shapes and tensor dimensions, I ask you to predict what will happen before I tell you, and I say plainly when a popular explanation is an oversimplification.

Knows about

  • convolutional networks
  • recurrent networks
  • transformers
  • attention
  • positional encodings
  • graph neural networks
  • reinforcement learning basics
  • debugging training
  • reading papers

Tutors by Nikolai

9 tutors

Attention mechanisms, step by step

Attention mechanisms, step by step

Compute attention by hand, then understand masks, heads, KV caching and efficient variantsIntermediateDeep learning4.3(4)66 lessonsSample
Nikolai Sorin$8
Debugging neural network training

Debugging neural network training

A systematic method for finding why a model will not train, diverges or quietly underperformsAdvancedDeep learning4.5(4)66 lessonsSample
Nikolai Sorin$10
The transformer, block by block

The transformer, block by block

Trace a token through every part of a transformer and count where the parameters liveAdvancedDeep learning4.7(3)60 lessonsSample
Nikolai Sorin$12
Convolutional networks from the pixel up

Convolutional networks from the pixel up

See how convolutions turn pixels into features, and calculate shapes and parameters yourselfBeginnerComputer vision4.7(3)56 lessonsSample
Nikolai Sorin$5
How to read a deep learning paper

How to read a deep learning paper

Read papers in passes, find the real claim and judge the evidence behind itAll levelsAI for research and study4.7(3)41 lessonsSample
Nikolai Sorin$6
Reinforcement learning basics

Reinforcement learning basics

Understand agents, rewards and policies, and compute a Q learning update yourselfBeginnerDeep learning4.7(3)41 lessonsSample
Nikolai SorinFree
RNNs, LSTMs and why transformers took over

RNNs, LSTMs and why transformers took over

Understand recurrent networks, their gates and limits, and the real reasons attention replaced themIntermediateDeep learning4.5(2)32 lessonsSample
Nikolai Sorin$7
Positional information in transformers

Positional information in transformers

Learn how transformers know word order, from sinusoids to rotary embeddings and long context limitsIntermediateDeep learningNew
Nikolai Sorin$8
Graph neural networks

Graph neural networks

Learn message passing on graphs and build models for nodes, edges and whole graphs without leakageAdvancedDeep learningNew
Nikolai Sorin$11

Recent reviews

What students said about Nikolai's tutors.

  • Sanna V.Sample

    Free and genuinely good. The reward hacking examples were funny and made me think about badly designed targets at my own job.

    On Reinforcement learning basics

  • Viktor H.Sample

    The 12d squared rule plus the embedding table let me sanity check a config file at work in about a minute. Lesson two on the residual stream changed how I picture the whole model.

    On The transformer, block by block

  • Leila N.Sample

    The tutor refused to guess and kept asking for my curves. Annoying at first, then I realised that was the lesson.

    On Debugging neural network training

  • Laura B.Sample

    The hand traced RNN in lesson one made backprop through time make sense. Before this, LSTM diagrams were just boxes and arrows to me.

    On RNNs, LSTMs and why transformers took over

  • Hyejin L.Sample

    I used to read papers front to back and give up by section three. The three pass method is simple and it works for me.

    On How to read a deep learning paper

  • Priya S.Sample

    The KV cache lesson explained why my local model runs out of memory with long chats. Estimating the cache size by hand was eye opening.

    On Attention mechanisms, step by step