Inference optimisation for serving at scale
Serve language models faster and cheaper by understanding prefill, decode, batching and caching
A taste of a lesson
Our streaming speed is fine with 5 users but drops sharply at 40. GPU compute shows only 50 percent busy. Why?
Compute being half idle suggests the limit is memory, not arithmetic. Two common causes. First, the KV cache: 40 users with long conversations may fill memory, so the scheduler runs smaller batches or evicts and recomputes. Second, long prompt prefills from new users may interrupt decoding for everyone. Check cache usage over time and how often requests wait or get preempted, and compare p95 inter token latency with and without long prompts in the mix. Prefix caching, a context cap or chunked prefill may help. Quick exercise: estimate per token cache size for your model using layers, KV heads and head dimension.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain why prefill is compute bound and decode is usually memory bandwidth bound
- Estimate KV cache memory and its effect on batch size and concurrency
- Choose between continuous batching, caching, speculative decoding and parallelism
- Measure latency with percentiles under realistic traffic
- Build a capacity plan with headroom and per request limits
Lesson plan
- 1 Prefill, decode and the metrics users feel Separate the two generation phases and link each to a latency metric. Start
- 2 Why decode is bandwidth bound Estimate decode speed limits from weight size and memory bandwidth. Start
- 3 KV cache management Size the KV cache and manage it so memory supports more concurrent users. Start
- 4 Continuous batching and scheduling Keep hardware busy without letting long requests hurt short ones. Start
- 5 Speculative decoding, parallelism and quantisation Pick the right advanced technique for a given bottleneck. Start
- 6 Load testing and capacity planning Test with realistic traffic and plan capacity with honest headroom. Start
Try asking
About this tutor
For engineers running model inference for many users who need to cut latency and cost without guessing. You will learn why generation has two very different phases, why decoding is usually limited by memory bandwidth rather than arithmetic, and how that explains the main optimisations. The lessons cover latency metrics that users feel, KV cache memory management, continuous batching, speculative decoding, splitting models across devices, and how quantisation interacts with all of these. You finish with capacity planning and load testing using realistic traffic shapes. Familiarity with transformers and basic serving concepts is assumed; the course stays neutral about specific serving frameworks.
Reviews
4.7
3 ratingsSample
- Viktor N.Sample
The bandwidth ceiling calculation changed how our team argues about hardware. We now estimate before we benchmark and the numbers line up surprisingly well.
- Tobias R.Sample
Load testing with real length distributions exposed that our synthetic tests were far too kind. Very practical course.
- Leila M.Sample
Dense but excellent on KV cache sizing and scheduling. I would have liked more on multi tenant fairness, though the tutor gave good pointers when asked.
About the teacher
Making models fast, small and affordable: hardware, quantisation, serving and edge
9 tutors 355 lessons taught Sample
I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...
See Magnus's profile and tutorsMore like this
Other tutors on the same or nearby topics.