Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Batching and caching for model inference

Serve more requests on the same hardware by batching smartly and caching what can be reused

By Magnus Eriksen Intermediate MLOps and deployment 4.7(3) 57 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $9 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Batching and caching for model inference AI tutor following Magnus Eriksen's plan
Student:

Our GPU utilisation looks low, yet we cannot add more concurrent users without running out of memory. Why?

Tutor:

That usually means the key value cache, not compute, is your limit. Each active request holds cache memory for every token in its context, and long prompts or long answers add up fast. Estimate it: 2 times layers times key value heads times head dimension times bytes per value gives memory per token; multiply by typical context length and concurrent users. Remedies include paged cache allocation to cut fragmentation, prefix caching for shared system prompts, a quantised cache, or capping context length. Quick exercise: with 128 KB per token and 3,000 token contexts, how many concurrent requests fit in 24 GB of free memory?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Distinguish latency, time to first token and throughput
  • Compare static, dynamic and continuous batching for different workloads
  • Estimate key value cache memory and its effect on concurrency
  • Apply prefix and response caching while managing staleness and privacy
  • Load test and read tail latency percentiles correctly

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Latency versus throughput Define the metrics that matter for inference and how batching trades them. Start
  2. 2 Static, dynamic and continuous batching Choose a batching strategy that fits your request pattern and model type. Start
  3. 3 KV cache memory Estimate cache size per request and see how it limits concurrency. Start
  4. 4 Prefix caching Reuse computation for shared prompt prefixes across requests. Start
  5. 5 Response and semantic caching Cache full answers where safe and avoid stale or leaked responses. Start
  6. 6 Queues, tail latency and load testing Predict and measure behaviour under realistic load. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers running model inference who want more throughput without letting latency get out of hand. You will separate latency from throughput, compare static, dynamic and continuous batching, and see why continuous batching matters so much for language models whose responses vary in length. You estimate key value cache memory per request, learn how paged cache allocation and prefix caching work in principle, and weigh response caching and semantic caching against staleness and privacy risks. The course finishes with queueing behaviour, tail latency and how to load test with realistic traffic. Examples stay tool neutral.

Reviews

4.7

3 ratingsSample

  • Maryam K.Sample

    The privacy point about response caches was important for us. We had been caching by question text only, across users.

  • Bruno T.Sample

    Load testing with real output length distributions changed our capacity plan completely. Our earlier tests used fixed short answers.

  • Sven L.Sample

    Estimating cache memory per request explained our concurrency ceiling in ten minutes. We moved the system prompt to the front and prefix caching started working.

About the teacher

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9 tutors 4.6(18) 355 lessons taught Sample

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...

See Magnus's profile and tutors