Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Deploying an open weights model

Choose, size, secure and run an open weights model in production, and compare its real cost

By Magnus Eriksen Intermediate MLOps and deployment 4.7(3) 46 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $9 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Deploying an open weights model AI tutor following Magnus Eriksen's plan
Student:

We want to self host an open model to save money compared with our API bill. How do I know if it will?

Tutor:

Compare total costs at your real traffic. On the API side, take last month's token usage and bill. On the self hosted side, estimate the hardware needed for peak concurrency, then its cost for every hour of the month, because idle GPUs still cost money. Add engineering time for setup, monitoring, updates and on call. Then check quality: run your own task set on the candidate model, since a cheaper model that needs more retries or human fixes is not cheaper. Self hosting tends to pay off with steady, high utilisation. Quick check: what share of the day does your traffic stay near its peak?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Shortlist open weights models using licence checks and your own task evaluations
  • Size hardware for weights, cache and expected concurrency
  • Secure a self hosted endpoint and verify model files
  • Pin versions, canary updates and keep a working rollback
  • Compare self hosting with hosted APIs using your real utilisation

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Licences and choosing a model Shortlist models whose licences fit your use and whose quality fits your task. Start
  2. 2 Serving software and APIs Understand what an inference server does and why standard APIs help. Start
  3. 3 Sizing hardware and quantised variants Estimate memory and throughput for your expected load. Start
  4. 4 Security for self hosted models Protect the endpoint, the model files and the systems the model touches. Start
  5. 5 Monitoring, updates and rollback Run the model reliably and change versions without surprises. Start
  6. 6 An honest cost comparison Compare self hosting with hosted APIs using your own traffic numbers. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers and technical leads who want to run an open weights model on their own infrastructure. You will learn how to read licence terms, shortlist models by evaluating them on your own task data, and size hardware for weights plus the context cache at your expected concurrency. The lessons cover inference servers in general terms, standard chat style APIs that let client code switch between models, quantised variants, authentication and rate limits, safe model file formats, and why prompt injection risks do not disappear when you self host. You finish with monitoring, version pinning and rollback, and an honest comparison of self hosting against hosted APIs.

Reviews

4.7

3 ratingsSample

  • Wei Z.Sample

    Pinning versions and canarying updates saved us when a newer model release changed its output format. We rolled back in minutes.

  • Gareth P.Sample

    Solid security lesson. I had not thought about weight formats that can execute code on load. Would have liked more on multi GPU serving.

  • Nadia L.Sample

    The cost lesson stopped us from a mistake. Our traffic is very bursty, and once we counted idle hours the API stayed cheaper for now.

About the teacher

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9 tutors 4.6(18) 355 lessons taught Sample

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...

See Magnus's profile and tutors