Initialisation and normalisation layers
Understand how weight initialisation and normalisation keep deep networks trainable
A taste of a lesson
Why does He initialisation use 2 over fan in instead of 1 over fan in?
Start from the goal: keep each layer's output variance about equal to its input variance. For a linear unit, output variance is roughly fan_in times weight variance times input variance, so weight variance 1/fan_in keeps it steady. But ReLU sets every negative pre-activation to zero, which for zero mean inputs removes about half of the second moment. Doubling the weight variance to 2/fan_in compensates for that halving, so the signal neither shrinks nor grows through many ReLU layers. Your turn: for a leaky ReLU with a small negative slope, would the ideal factor be a bit above or a bit below 2?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Derive why activation variance should stay stable across layers
- Choose Xavier, He or scaled initialisation for a given architecture
- Compare batch, layer, RMS and group normalisation and their failure modes
- Explain why pre norm transformers are easier to train at depth
- Diagnose instability caused by initialisation or normalisation choices
Lesson plan
- 1 Why initialisation matters Show how weight scale controls whether signals vanish or explode through depth. Start
- 2 Xavier, He and depth aware schemes Match initialisation variance to the activation function and architecture. Start
- 3 Batch normalisation in depth Understand what batch norm computes and where it causes problems. Start
- 4 Layer norm, RMSNorm and group norm Compare per example normalisation methods and pick one for a model. Start
- 5 Normalisation placement in transformers Explain the stability difference between pre norm and post norm designs. Start
- 6 Diagnosing instability Use activation and gradient statistics to locate initialisation and normalisation problems. Start
Try asking
About this tutor
For practitioners who train deep networks and want to understand the settings that quietly decide whether training is stable. You will derive why activation variance must stay roughly constant across layers, compare zero, Xavier and He initialisation, and look at how residual networks and transformers scale their initial weights. The second half covers batch normalisation, layer normalisation, RMSNorm and group normalisation: what each normalises over, how they behave at training and inference time, and why pre norm transformers train more stably than post norm ones. Expect light derivations, small numerical experiments to reason about, and honest notes on which explanations are still debated.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
I teach how neural networks learn, one small worked example at a time
9 tutors 424 lessons taught Sample
I teach the core mechanics of deep learning: what a neuron computes, how a loss turns mistakes into numbers, and how gradients and optimisers change weights. My background is in building and training models for applied research teams, which mostly meant staring at loss curves that refused to go down. That shaped how I teach. I start every idea with...
See Mira's profile and tutorsMore like this
Other tutors on the same or nearby topics.