Tokens: how models read text
See how text becomes tokens and why that explains odd model behaviour, limits and costs
A taste of a lesson
Why did the chatbot say 'strawberry' has two r's?
Because it never saw the letters directly. Before the model reads 'strawberry', a tokenizer splits it into chunks, perhaps something like 'str', 'aw' and 'berry', depending on the model. The model sees ID numbers for those chunks, not individual letters, so counting r's means reasoning about what is inside each chunk, which it learned only indirectly. Newer models often get it right, especially if they spell the word out first. Try asking it to spell the word letter by letter before counting. Does the answer change?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what a token is and why models use word pieces
- Estimate roughly how many tokens a piece of English text contains
- Explain how tokens cause letter counting and arithmetic mistakes
- Describe why token counts drive context limits and costs, and vary by language
Lesson plan
- 1 From text to pieces Understand what tokens are and see text split into them. Start
- 2 How a tokenizer is built Learn the basic idea of building a vocabulary by merging frequent pieces. Start
- 3 Token effects you can see Explain everyday model quirks using tokens. Start
- 4 Tokens across languages and code See why token efficiency differs between languages and formats. Start
- 5 Estimating and budgeting tokens Use rough rules to estimate token counts for limits and costs. Start
Try asking
About this tutor
For beginners and curious users who want one concrete piece of how language models work. You learn what tokens are, why models split text into word pieces rather than letters or whole words, and how a tokenizer is built. Then you use that idea to explain everyday puzzles: why models miscount letters, struggle with some arithmetic, handle some languages less efficiently, and why usage limits and prices are counted in tokens. You practise estimating token counts and spotting token effects in real answers. No coding is required, though you can try a public tokenizer viewer if you like.
Reviews
4.7
3 ratingsSample
- Bilal H.Sample
Short, precise and honest about approximations. I now explain tokens to my colleagues with the strawberry example.
- Yuki M.Sample
Great hands on examples. The tokenizer building lesson was a little abstract, but the spelling trick at the end really worked.
- Daniel F.Sample
Finally understand why my API bill is counted in tokens. The tip about other languages costing more explained why our Polish support replies were pricier.
About the teacher
I explain how language models really work, from tokens to attention, without hand waving
9 tutors 525 lessons taught Sample
I like taking the mystery out of language models. I teach what happens between typing a question and reading an answer: tokens, context windows, embeddings, attention, training and fine tuning. I start every topic with a picture or a small worked example and only add maths when it earns its place. My work background is in software and teaching workshops...
See Bastian's profile and tutorsMore like this
Other tutors on the same or nearby topics.