Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Diffusion models from noise to sample

Understand how diffusion models learn to remove noise and how guidance and latents shape the result

By Mateo Rojas Advanced Deep learning Image and video AI 4.7(3) 44 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Diffusion models from noise to sample AI tutor following Mateo Rojas's plan
Student:

Why does raising the guidance scale make images match the prompt better but look worse at high values?

Tutor:

Guidance computes two noise predictions, one with the prompt and one without, and moves further along their difference: uncond + w times (cond minus uncond). That difference points towards features the prompt makes more likely, so a larger w pushes harder towards them. But it also extrapolates beyond anything the model saw in training, so colours and contrast get exaggerated and samples become similar to each other. Moderate values balance prompt adherence against naturalness and variety. Quick check: what does a guidance scale of exactly 1 reduce to?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Derive the forward noising process and the noise prediction training objective
  • Explain how sampling turns noise into data and why fewer step samplers work
  • Describe latent diffusion and how text conditioning enters the model
  • Use classifier free guidance knowingly and explain its trade offs
  • Assess diffusion evaluation metrics and compute costs realistically

Lesson plan

6 lessons. Pick one to start there.

  1. 1 The forward process Describe how data is gradually turned into noise and why any step can be sampled directly. Start
  2. 2 Training to predict noise Understand the simple mean squared error objective and its variants. Start
  3. 3 Sampling and faster samplers Run the reverse process and see how step count trades against quality. Start
  4. 4 Latent diffusion and backbones Explain why generation in an autoencoder's latent space is cheaper and what backbones are used. Start
  5. 5 Conditioning and guidance Control generation with prompts and understand classifier free guidance. Start
  6. 6 Evaluation and cost Judge diffusion results and estimate what training and sampling really cost. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who understand neural network training and want a solid grasp of how modern image, audio and video generators work under the hood. You will derive the forward process that gradually adds noise, the simple training objective of predicting that noise, and the reverse process that turns pure noise into a sample. Then you study the practical engineering: noise schedules, samplers that need far fewer steps, latent diffusion with an autoencoder, conditioning on text through cross attention, classifier free guidance and its trade offs, and the move from U-Net to transformer backbones. The course ends with evaluation caveats and the real compute costs of training and sampling.

Reviews

4.7

3 ratingsSample

  • Brigid M.Sample

    Rigorous and well paced. The sampler lesson could use a bit more on why higher order solvers work, but the practical trade offs were clear.

  • Noemi F.Sample

    The point that predicting noise, the clean image and velocity are reparameterisations of each other cleared up weeks of confusion from reading different papers.

  • Kaito H.Sample

    Excellent on guidance. I now pick guidance scales deliberately instead of copying values from forums.

About the teacher

Mateo Rojas

Generative models and representation learning, explained with intuition first and maths second

9 tutors 4.6(20) 335 lessons taught Sample

I teach how models learn useful representations and how they generate new data: autoencoders, GANs, diffusion models, self supervised learning and language model pretraining. I came to this through research engineering work where we had to decide which kind of model was worth the compute, so I teach with trade offs in mind. Each topic starts with a picture or...

See Mateo's profile and tutors