Tokenisation: how text becomes numbers
See how models split text into tokens, why it matters for cost and context, and where it trips up
A taste of a lesson
Why does the same paragraph cost more when I send it in Georgian than in English?
Because the tokeniser learned its pieces mostly from text in languages that were common in its training data. English words often get one token each, while words in a script that appeared less often are split into many small pieces, sometimes single characters or even bytes. More tokens means a higher price, slower responses and less room in the context window for the same meaning. Newer tokenisers have improved this for many languages, so it varies by model. Quick check: if a word becomes five tokens instead of one, what happens to how much text fits in the same context window?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain why models split text into tokens and what a token ID is
- Run a small byte pair encoding example by hand
- Explain why token counts differ across languages and tokenisers
- Predict quirks with spelling, letter counting and long numbers
Lesson plan
- 1 Why text must become numbers Understand tokens, token IDs and embeddings as the first step in any language model. Start
- 2 Characters, words and subwords Compare splitting strategies and see why subwords became the standard. Start
- 3 Byte pair encoding by hand Learn a small vocabulary from a toy corpus using repeated pair merges. Start
- 4 Tokens across languages See why the same message can use very different numbers of tokens in different languages. Start
- 5 Quirks and practical tips Explain odd model behaviour with spelling and numbers and count tokens sensibly. Start
Try asking
About this tutor
For beginners who want to understand the first step every language model takes: turning text into tokens. You will compare splitting text into characters, words and subword pieces, then run a small byte pair encoding example by hand to see how a vocabulary is learned from data. The lessons explain why the same sentence costs different numbers of tokens in different languages and with different tokenisers, how that affects price and how much fits in a context window, and why models can stumble on spelling, counting letters and long numbers. No coding is needed, though learners who code get tips for counting tokens in their own projects.
Reviews
4.7
3 ratingsSample
- Akosua D.Sample
Good for beginners. I would have liked one more example with numbers and dates, but the main ideas were well explained.
- Oscar P.Sample
The letter counting explanation made a lot of strange chatbot mistakes make sense. Short course, very clear.
- Nino B.Sample
Finally I understand why my Georgian prompts were so expensive. The BPE by hand lesson was surprisingly fun.
About the teacher
Practical NLP: from tokens and embeddings to classification, translation and speech
9 tutors 470 lessons taught Sample
I teach natural language processing as a craft: turning messy text in many languages into something a model can use, and checking honestly whether the result works. I grew up switching between Arabic, French and English, and my work has been on text and speech systems that had to serve speakers of more than one language, so I notice quickly...
See Nadia's profile and tutorsMore like this
Other tutors on the same or nearby topics.