Quantisation: smaller, faster models
Shrink models with lower precision numbers and measure exactly what quality you trade away
A taste of a lesson
I quantised my model to 4 bit and it uses far less memory, but it is slower than the 16 bit version. How?
Memory savings follow directly from the arithmetic of bytes, but speed depends on how the computation runs. If your runtime stores weights in 4 bit and converts them back to 16 bit with a separate step before each multiplication, that extra work can outweigh the bandwidth saved. Fast setups use kernels that fuse dequantisation into the matrix multiplication. Also check batch size: at large batches you are compute bound, and weight only quantisation helps less. Benchmark both models at your real batch size and sequence length. Quick check: at batch size 1, is generation usually limited by arithmetic or by memory bandwidth?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compare common number formats and their memory costs
- Quantise and dequantise values by hand using scale and zero point
- Explain how outliers and scale granularity affect quality
- Choose between weight only, weight and activation, PTQ and QAT approaches
- Measure quality loss and speed on your own task and hardware
Lesson plan
- 1 Number formats and memory Compare 32, 16, 8 and 4 bit formats and compute model sizes for each. Start
- 2 Scale and zero point by hand Map real values to integers and back, and measure the rounding error. Start
- 3 Granularity and outliers Understand why a few large values make low bit quantisation hard. Start
- 4 What to quantise and when Choose between weight only and weight and activation quantisation, and PTQ or QAT. Start
- 5 Measuring quality loss Evaluate a quantised model on the tasks that matter to you. Start
- 6 Speed, kernels and the cache Measure real speed and extend savings to the key value cache. Start
Try asking
About this tutor
For practitioners who want to run models with less memory and often more speed by storing numbers in lower precision. You will learn the common number formats, compute a quantised value and its error by hand with a scale and zero point, and see why outliers make low bit quantisation hard. The lessons cover per tensor, per channel and per group scales, weight only versus weight and activation quantisation, post training methods with calibration data versus quantisation aware training, and quantising the key value cache. Most importantly, you learn to measure quality loss on your own task, and why smaller does not always mean faster.
Reviews
4.3
3 ratingsSample
- Ravi S.Sample
Doing the scale and zero point example by hand made every later lesson easier. I now understand why group size appears in quantisation settings.
- Kofi A.Sample
Good balance of theory and practice. I would have liked more on QAT, but the tutor was clear it is less common for most deployments.
- Elena V.Sample
The point about testing maths and other languages separately was spot on. Our 4 bit model was fine in English and clearly worse in Polish.
About the teacher
Making models fast, small and affordable: hardware, quantisation, serving and edge
9 tutors 355 lessons taught Sample
I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...
See Magnus's profile and tutorsMore like this
Other tutors on the same or nearby topics.