Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Tokens: how models read text

See how text becomes tokens and why that explains odd model behaviour, limits and costs

By Bastian Weber Beginner How language models work 4.7(3) 54 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $4 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Tokens: how models read text AI tutor following Bastian Weber's plan
Student:

Why did the chatbot say 'strawberry' has two r's?

Tutor:

Because it never saw the letters directly. Before the model reads 'strawberry', a tokenizer splits it into chunks, perhaps something like 'str', 'aw' and 'berry', depending on the model. The model sees ID numbers for those chunks, not individual letters, so counting r's means reasoning about what is inside each chunk, which it learned only indirectly. Newer models often get it right, especially if they spell the word out first. Try asking it to spell the word letter by letter before counting. Does the answer change?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain what a token is and why models use word pieces
  • Estimate roughly how many tokens a piece of English text contains
  • Explain how tokens cause letter counting and arithmetic mistakes
  • Describe why token counts drive context limits and costs, and vary by language

Lesson plan

5 lessons. Pick one to start there.

  1. 1 From text to pieces Understand what tokens are and see text split into them. Start
  2. 2 How a tokenizer is built Learn the basic idea of building a vocabulary by merging frequent pieces. Start
  3. 3 Token effects you can see Explain everyday model quirks using tokens. Start
  4. 4 Tokens across languages and code See why token efficiency differs between languages and formats. Start
  5. 5 Estimating and budgeting tokens Use rough rules to estimate token counts for limits and costs. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For beginners and curious users who want one concrete piece of how language models work. You learn what tokens are, why models split text into word pieces rather than letters or whole words, and how a tokenizer is built. Then you use that idea to explain everyday puzzles: why models miscount letters, struggle with some arithmetic, handle some languages less efficiently, and why usage limits and prices are counted in tokens. You practise estimating token counts and spotting token effects in real answers. No coding is required, though you can try a public tokenizer viewer if you like.

Reviews

4.7

3 ratingsSample

  • Bilal H.Sample

    Short, precise and honest about approximations. I now explain tokens to my colleagues with the strawberry example.

  • Yuki M.Sample

    Great hands on examples. The tokenizer building lesson was a little abstract, but the spelling trick at the end really worked.

  • Daniel F.Sample

    Finally understand why my API bill is counted in tokens. The tip about other languages costing more explained why our Polish support replies were pricier.

About the teacher

Bastian Weber

I explain how language models really work, from tokens to attention, without hand waving

9 tutors 4.5(24) 525 lessons taught Sample

I like taking the mystery out of language models. I teach what happens between typing a question and reading an answer: tokens, context windows, embeddings, attention, training and fine tuning. I start every topic with a picture or a small worked example and only add maths when it earns its place. My work background is in software and teaching workshops...

See Bastian's profile and tutors