Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.

Teacher since October 2025

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9

tutors built

4.6Sample

average from 18 reviews

355Sample

lessons taught by their tutors

About Magnus

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to explain things in terms of bytes, bandwidth and latency budgets. I like back of the envelope estimates that you can do before spending money, and I am careful to separate durable principles from details that change with every new chip or library release.

Knows about

  • GPUs and accelerators
  • training cost estimation
  • quantisation
  • distillation
  • batching and caching
  • serving optimisation
  • deploying open weights models
  • edge and on device models
  • visual inspection

Tutors by Magnus

9 tutors

Inference optimisation for serving at scale

Inference optimisation for serving at scale

Serve language models faster and cheaper by understanding prefill, decode, batching and cachingAdvancedMLOps and deployment4.7(3)83 lessonsSample
Magnus Eriksen$15
Planning the cost of a training run

Planning the cost of a training run

Estimate compute, time, memory and budget for a training or fine tuning run before you spendAll levelsFine tuning and training4.3(3)65 lessonsSample
Magnus Eriksen$7
Batching and caching for model inference

Batching and caching for model inference

Serve more requests on the same hardware by batching smartly and caching what can be reusedIntermediateMLOps and deployment4.7(3)57 lessonsSample
Magnus Eriksen$9
GPUs and AI hardware for beginners

GPUs and AI hardware for beginners

Understand what GPUs do for AI and estimate whether a model will fit on your hardwareBeginnerMLOps and deployment4.7(3)55 lessonsSample
Magnus EriksenFree
Quantisation: smaller, faster models

Quantisation: smaller, faster models

Shrink models with lower precision numbers and measure exactly what quality you trade awayIntermediateFine tuning and training4.3(3)49 lessonsSample
Magnus Eriksen$9
Deploying an open weights model

Deploying an open weights model

Choose, size, secure and run an open weights model in production, and compare its real costIntermediateMLOps and deployment4.7(3)46 lessonsSample
Magnus Eriksen$9
AI visual inspection for factories, explained

AI visual inspection for factories, explained

Learn how camera based defect detection works on a production line, and what makes it succeedBeginnerComputer visionNew
Magnus Eriksen$5
On device and edge models

On device and edge models

Decide when a model should run on the phone or device itself, and make it fit and run wellBeginnerMLOps and deploymentNew
Magnus Eriksen$6
Knowledge distillation: teaching a smaller model

Knowledge distillation: teaching a smaller model

Train compact student models from large teachers with soft targets, generated data and careful evaluationAdvancedDeep learningNew
Magnus Eriksen$12

Recent reviews

What students said about Magnus's tutors.

  • Viktor N.Sample

    The bandwidth ceiling calculation changed how our team argues about hardware. We now estimate before we benchmark and the numbers line up surprisingly well.

    On Inference optimisation for serving at scale

  • Wei Z.Sample

    Pinning versions and canarying updates saved us when a newer model release changed its output format. We rolled back in minutes.

    On Deploying an open weights model

  • Tobias R.Sample

    Load testing with real length distributions exposed that our synthetic tests were far too kind. Very practical course.

    On Inference optimisation for serving at scale

  • Patrick D.Sample

    I presented a one page estimate with all assumptions listed and our finance lead approved it the same day. The buffer for failed runs turned out to be needed.

    On Planning the cost of a training run

  • Maryam K.Sample

    The privacy point about response caches was important for us. We had been caching by question text only, across users.

    On Batching and caching for model inference

  • Bruno T.Sample

    Load testing with real output length distributions changed our capacity plan completely. Our earlier tests used fixed short answers.

    On Batching and caching for model inference