Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Inference optimisation for serving at scale

Serve language models faster and cheaper by understanding prefill, decode, batching and caching

By Magnus Eriksen Advanced MLOps and deployment 4.7(3) 83 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $15 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Inference optimisation for serving at scale AI tutor following Magnus Eriksen's plan
Student:

Our streaming speed is fine with 5 users but drops sharply at 40. GPU compute shows only 50 percent busy. Why?

Tutor:

Compute being half idle suggests the limit is memory, not arithmetic. Two common causes. First, the KV cache: 40 users with long conversations may fill memory, so the scheduler runs smaller batches or evicts and recomputes. Second, long prompt prefills from new users may interrupt decoding for everyone. Check cache usage over time and how often requests wait or get preempted, and compare p95 inter token latency with and without long prompts in the mix. Prefix caching, a context cap or chunked prefill may help. Quick exercise: estimate per token cache size for your model using layers, KV heads and head dimension.

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain why prefill is compute bound and decode is usually memory bandwidth bound
  • Estimate KV cache memory and its effect on batch size and concurrency
  • Choose between continuous batching, caching, speculative decoding and parallelism
  • Measure latency with percentiles under realistic traffic
  • Build a capacity plan with headroom and per request limits

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Prefill, decode and the metrics users feel Separate the two generation phases and link each to a latency metric. Start
  2. 2 Why decode is bandwidth bound Estimate decode speed limits from weight size and memory bandwidth. Start
  3. 3 KV cache management Size the KV cache and manage it so memory supports more concurrent users. Start
  4. 4 Continuous batching and scheduling Keep hardware busy without letting long requests hurt short ones. Start
  5. 5 Speculative decoding, parallelism and quantisation Pick the right advanced technique for a given bottleneck. Start
  6. 6 Load testing and capacity planning Test with realistic traffic and plan capacity with honest headroom. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers running model inference for many users who need to cut latency and cost without guessing. You will learn why generation has two very different phases, why decoding is usually limited by memory bandwidth rather than arithmetic, and how that explains the main optimisations. The lessons cover latency metrics that users feel, KV cache memory management, continuous batching, speculative decoding, splitting models across devices, and how quantisation interacts with all of these. You finish with capacity planning and load testing using realistic traffic shapes. Familiarity with transformers and basic serving concepts is assumed; the course stays neutral about specific serving frameworks.

Reviews

4.7

3 ratingsSample

  • Viktor N.Sample

    The bandwidth ceiling calculation changed how our team argues about hardware. We now estimate before we benchmark and the numbers line up surprisingly well.

  • Tobias R.Sample

    Load testing with real length distributions exposed that our synthetic tests were far too kind. Very practical course.

  • Leila M.Sample

    Dense but excellent on KV cache sizing and scheduling. I would have liked more on multi tenant fairness, though the tutor gave good pointers when asked.

About the teacher

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9 tutors 4.6(18) 355 lessons taught Sample

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...

See Magnus's profile and tutors