On device and edge models
Decide when a model should run on the phone or device itself, and make it fit and run well
A taste of a lesson
I want to run a 3 billion parameter language model in my phone app. Is that realistic?
Let us estimate. At 16 bit, 3 billion parameters need about 6 GB just for weights, which is more than most phones can give one app. At 4 bit it drops to roughly 1.5 GB, plus extra for activations and the KV cache, so it may fit on recent high end phones but not older ones. Then check speed, battery and heat during a long chat, not one prompt. A smaller model for common requests, with the cloud as fallback, is often the safer design. Quick exercise: estimate the weight size of a 1 billion parameter model at 8 bit.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Decide whether a use case benefits from on device inference or belongs in the cloud
- Estimate a model's memory footprint from parameter count and precision
- Explain how quantisation, pruning and small architectures shrink models
- Benchmark a model on real target hardware including heat and battery
- Design a hybrid device and cloud setup with safe model updates
Lesson plan
- 1 Why run on the device at all Weigh latency, privacy, offline use and cost against the extra work of on device models. Start
- 2 Budgets: memory, compute, power and heat Estimate whether a model fits a device's memory and power budget. Start
- 3 Making models smaller Compare small architectures, quantisation, pruning and distillation for edge use. Start
- 4 Converting for an on device runtime Export a trained model and confirm it runs correctly on the target runtime. Start
- 5 Benchmarking on real hardware Measure what users will actually experience, not desktop numbers. Start
- 6 Updates and hybrid designs Ship, update and roll back models safely, and split work between device and cloud. Start
Try asking
About this tutor
For developers, product people and students who want to run models on phones, laptops, cameras or small boards instead of in the cloud. You will learn when on device inference is worth it (latency, privacy, offline use, running cost) and what it costs you in memory, power, heat and update complexity. The lessons cover small architectures, quantisation and pruning at a practical level, converting a trained model to an on device runtime in general terms, and benchmarking on the real target hardware rather than your desktop. You finish by designing a hybrid setup where the device handles common cases and the cloud handles the rest. No hardware background is needed.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Making models fast, small and affordable: hardware, quantisation, serving and edge
9 tutors 355 lessons taught Sample
I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...
See Magnus's profile and tutorsMore like this
Other tutors on the same or nearby topics.