Named entity recognition in practice
Extract people, places, organisations and custom entities from text, and evaluate the results properly
A taste of a lesson
My NER model has 97 percent token accuracy but extracted company names are often cut off. Why the gap?
Token accuracy is dominated by O tokens, the words that are not entities, so a model can score very high while still getting entity boundaries wrong. Switch to entity level evaluation with strict matching: a company name predicted as 'Global Freight' when the gold span is 'Global Freight Ltd' counts as a miss. Then look at the errors: are suffixes like Ltd inconsistently labelled in your guidelines, or are long names split into many subwords? Quick exercise: under strict scoring, how many false positives and false negatives does that one truncated prediction create?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Tag text with the BIO scheme and align labels with subword tokens
- Write annotation guidelines that settle boundary and nesting decisions
- Evaluate NER at entity level with strict and partial matching
- Choose between a fine tuned tagger and LLM extraction for a project
- Explain entity linking and the limits of automated personal data detection
Lesson plan
- 1 Entities and the BIO scheme Represent entity spans as token labels and tag example sentences by hand. Start
- 2 Guidelines and hard boundaries Write rules for titles, abbreviations, nesting and other boundary cases. Start
- 3 Subword alignment in transformer taggers Map word level labels to subword tokens and back without losing entities. Start
- 4 Entity level evaluation Score predictions at entity level with strict and partial matching. Start
- 5 Fine tuned taggers versus LLM extraction Compare approaches on accuracy, cost, speed and maintenance for your case. Start
- 6 Linking entities and handling personal data Connect mentions to real records and understand the limits of PII detection. Start
Try asking
About this tutor
For learners with basic NLP knowledge who need to pull structured information out of text: names, organisations, products, dates, drug names or any custom entity type. You will learn BIO tagging, how transformer taggers deal with subword tokens, and why nested and overlapping entities complicate things. You will write annotation guidelines, compare entity level evaluation in strict and partial modes, and decide between fine tuning a tagger and prompting a large language model for extraction. The course also covers linking entities to a knowledge base and the limits of automatic detection of personal data, with examples in several languages and scripts.
Reviews
4.7
3 ratingsSample
- Ravi S.Sample
The span verification tip for LLM extraction was gold. Our model was quietly rewriting drug names.
- Katarzyna W.Sample
Detailed and precise. The subword alignment lesson solved a bug where Polish names lost their endings.
- Yaw A.Sample
Switching to entity level strict scoring dropped our headline number a lot, but finally showed the real problem with suffixes.
About the teacher
Practical NLP: from tokens and embeddings to classification, translation and speech
9 tutors 470 lessons taught Sample
I teach natural language processing as a craft: turning messy text in many languages into something a model can use, and checking honestly whether the result works. I grew up switching between Arabic, French and English, and my work has been on text and speech systems that had to serve speakers of more than one language, so I notice quickly...
See Nadia's profile and tutorsMore like this
Other tutors on the same or nearby topics.