Convolutional networks from the pixel up
See how convolutions turn pixels into features, and calculate shapes and parameters yourself
A taste of a lesson
A 64 by 64 input with 3 channels goes through a 5 by 5 conv, 16 filters, stride 1, no padding. What comes out?
Use the formula per spatial dimension: (64 - 5 + 0) / 1 + 1 = 60. So each feature map is 60 by 60, and with 16 filters you get 16 channels: an output of 60 by 60 by 16. Parameters: each filter spans 5 by 5 by 3 = 75 weights plus one bias, so 76 per filter, times 16 gives 1,216. Notice how small that is compared with a dense layer on the same input. Now try it yourself: what padding would keep the output at 64 by 64 with the same kernel?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute a small convolution by hand and explain what a feature map is
- Calculate output shapes from kernel size, stride and padding
- Count the parameters of convolutional layers and whole small networks
- Describe how residual connections and pooling shape modern CNNs
Lesson plan
- 1 Images as numbers Represent images as tensors and see why dense layers handle them poorly. Start
- 2 The convolution operation by hand Slide a 3 by 3 kernel over a small grid and compute the feature map. Start
- 3 Channels, stride and padding Use the output size formula and handle multi channel inputs and outputs. Start
- 4 Pooling and receptive fields Understand how networks grow their view of the image as they go deeper. Start
- 5 A small classifier, end to end Assemble a CNN classifier and trace shapes from input to class scores. Start
- 6 Residual connections and honest limits See why residual designs dominate and where CNNs still fall short. Start
Try asking
About this tutor
For beginners who know what a basic neural network is and want to understand how networks see images. You start with an image as a grid of numbers, slide a small filter across it by hand, and watch an edge detector appear. From there you build up feature maps, channels, stride, padding and pooling, then stack layers into a small classifier and calculate its output shapes and parameter count. You finish with the ideas behind modern designs, such as residual connections and global pooling, and an honest look at what convolutional networks do and do not understand about images. Light maths, lots of small diagrams described in words.
Reviews
4.7
3 ratingsSample
- Rahul V.Sample
Computing the edge detector by hand made feature maps feel real. I finally understand what people mean by weight sharing.
- Emma J.Sample
The shape tracing lesson was excellent. I would have liked a slightly longer section on residual blocks, but it covered the main idea.
- Kwame O.Sample
I liked the honesty about texture bias. It changed how I think about why my model failed on sketches.
About the teacher
Architectures explained from the inside: convolutions, recurrence, attention and beyond
9 tutors 362 lessons taught Sample
I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...
See Nikolai's profile and tutorsMore like this
Other tutors on the same or nearby topics.