Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Reinforcement learning basics

Understand agents, rewards and policies, and compute a Q learning update yourself

By Nikolai Sorin Beginner Deep learning Machine learning 4.7(3) 41 lessons taught Sample

This tutor is free. Your lesson starts right away.

A taste of a lesson

One question and the answer, as this tutor gives it.

Reinforcement learning basics AI tutor following Nikolai Sorin's plan
Student:

In the Q learning update, why not just replace the old value with the new target?

Tutor:

Because one experience can be misleading. Rewards and next states are often random, so a single lucky or unlucky step would throw the estimate around. The learning rate alpha blends old and new: with alpha 0.5 you move halfway towards the target each time, so the value settles towards an average over many experiences. Setting alpha to 1 means full replacement, which only works when everything is predictable. Tiny exercise: your Q value is 3, the target is 5 and alpha is 0.2. What is the new Q value?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Describe agent, environment, state, action, reward, policy and value in plain words
  • Compute discounted returns and a Q learning update by hand
  • Explain the exploration versus exploitation trade off using epsilon greedy
  • Recognise reward hacking and why reward design is difficult
  • Explain in broad terms how RL ideas are used to tune language models

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Agents, environments and rewards Name the parts of an RL problem using a small grid world. Start
  2. 2 Returns and discounting Compute discounted returns and see how the discount shapes behaviour. Start
  3. 3 Exploration versus exploitation Balance trying new actions against using the best known one in a bandit problem. Start
  4. 4 Values and Q learning Estimate action values and update them from experience with the Q learning rule. Start
  5. 5 Policies learned directly Understand policy gradient and actor critic methods in plain words. Start
  6. 6 Reward design and feedback from people See why rewards get gamed and how preference feedback is used for language models. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For beginners who want to know how machines learn from trial and error rather than from labelled examples. You will meet the core ideas through small games: an agent, an environment, states, actions and rewards, and the goal of maximising total reward over time. You will compute discounted returns, see the exploration versus exploitation dilemma in a slot machine problem, and work through a Q learning update by hand. Later lessons explain policy gradient methods in words, why reward design is hard, and how reinforcement learning ideas appear when language models are tuned with human feedback. Light maths, all of it explained step by step.

Reviews

4.7

3 ratingsSample

  • Sanna V.Sample

    Free and genuinely good. The reward hacking examples were funny and made me think about badly designed targets at my own job.

  • Lucia P.Sample

    The slot machine lesson made exploration click instantly. I kept picking the first machine that paid out, which was exactly the point.

  • Ahmed R.Sample

    Clear and friendly. The Q learning numbers were helpful. I wanted more on how chatbots use this, but the tutor was honest that it is only a light overview here.

About the teacher

Nikolai Sorin

Architectures explained from the inside: convolutions, recurrence, attention and beyond

9 tutors 4.6(22) 362 lessons taught Sample

I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...

See Nikolai's profile and tutors