Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Knowledge distillation: teaching a smaller model

Train compact student models from large teachers with soft targets, generated data and careful evaluation

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Knowledge distillation: teaching a smaller model AI tutor following Magnus Eriksen's plan
Student:

My distilled student matches the teacher's average accuracy within 2 points. Can I ship it?

Tutor:

Not on that number alone. Averages hide where students usually lose: rare classes, long or unusual inputs, and behaviour the teacher handled carefully, such as refusing unsafe requests. Break the evaluation into slices: frequency of the class or intent, input length, languages, and your known hard cases. Compare student and teacher per slice and look at examples where they disagree. Also confirm latency and cost gains on real hardware, since that is the point of distilling. Quick exercise: name two slices in your data where you would expect the student to fall furthest behind, and why.

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Compute temperature softened probabilities and explain what they reveal
  • Build a distillation loss with correct temperature scaling and weighting
  • Run sequence level distillation for a language model with filtered teacher outputs
  • Choose a student size with the capacity gap in mind
  • Evaluate a distilled student by slice and check licence terms before training

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Soft targets and temperature Compute softened teacher probabilities and see what information they carry. Start
  2. 2 The distillation loss Combine soft target KL and hard label cross entropy with the right scaling. Start
  3. 3 Feature and intermediate layer distillation Match hidden representations when output matching alone is not enough. Start
  4. 4 Distilling language models Fine tune a student on teacher generated responses and filter them well. Start
  5. 5 Student size and the capacity gap Choose a student that can absorb the teacher's behaviour for your task. Start
  6. 6 Evaluation, licences and copied flaws Judge the student by slice and respect the terms of the teacher model. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who need a smaller, cheaper model that keeps as much of a large model's quality as possible. You will work through classic distillation with soft targets and temperature, compute the softened probabilities by hand, and build the combined loss with the right scaling. Then you move to the methods used with language models today: training on teacher generated outputs, feature and intermediate layer distillation, and choosing a student size that can actually absorb what the teacher knows. The final lessons cover licence and terms of use limits on training from another model's outputs, and how to evaluate a student honestly beyond its average score.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9 tutors 4.6(18) 355 lessons taught Sample

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...

See Magnus's profile and tutors