AI for Archivists: OCR, Description and Ethics
Use OCR, handwriting recognition and AI description to open up collections while respecting context
A taste of a lesson
Our handwriting recognition gets about 85% of words right on 19th century letters. Is that usable?
It depends on the purpose. For keyword search, 85% word accuracy can be quite useful, especially with fuzzy matching, because most distinctive words will be found. For publishing transcriptions, it's not enough: roughly one word in seven is wrong, and errors cluster in names and places, exactly what researchers search for. A common approach: publish raw output labelled as machine generated for search, and correct names, dates and places first. Which use matters most to your users?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain how OCR and handwriting recognition work and what affects their quality
- Measure transcription quality and decide what level of correction is needed
- Draft descriptions and index terms with AI and review them against archival standards
- Address sensitive information, harmful language and transparency in AI generated metadata
Lesson plan
- 1 OCR and HTR basics Understand how text recognition works and what makes it succeed or fail. Start
- 2 Measuring and improving quality Assess error rates and choose the right level of correction for each purpose. Start
- 3 AI assisted description Draft summaries and index terms with AI and review them against standards. Start
- 4 Sensitive records and access Protect privacy and consider harm before processing or publishing records. Start
- 5 Harmful language and transparency Handle historical language responsibly and label machine generated metadata. Start
Try asking
About this tutor
For archivists, records managers, special collections staff and digitisation volunteers. We cover how optical character recognition and handwritten text recognition work, how to judge and improve transcription quality, using AI to draft descriptions, summaries and index terms, and searching large digitised collections. We take ethics seriously: provenance and original order, sensitive and personal information in records, harmful historical language, bias in AI generated description, and transparency about machine generated metadata. Practical, collection focused and respectful of archival principles.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Translator, editor and digitisation hand teaching AI for words, archives, buildings and travel
9 tutors 307 lessons taught Sample
I began as a technical translator for engineering and architecture firms, worked on a regional news desk as a sub editor, and later joined digitisation projects with libraries and archives. Translating for tourism businesses brought me into hospitality too. What links all of it is careful work with text and documents where a small error travels a long way. I...
See Zoltan's profile and tutorsMore like this
Other tutors on the same or nearby topics.