Summarisation systems and their failure modes
Build and judge summaries that stay faithful to the source, from short notes to long reports
A taste of a lesson
My summary of a clinical study says the drug 'reduces symptoms', but the paper says 'may reduce'. Is that a big deal?
Yes, it is one of the most important failure types. 'May reduce' signals uncertainty, perhaps a small sample or a result that was not statistically clear; 'reduces' turns that into a firm claim. A reader acting on the summary would be misled even though every word sounds reasonable. Fixes: ask the summariser to preserve hedging and limitations, add a line for study size and caveats, and check claims against the source. Quick exercise: rewrite the sentence so it is short but keeps the original level of certainty.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain the difference between extractive and abstractive summaries and their typical risks
- Spot unsupported details, dropped caveats and misplaced emphasis in a summary
- Design a pipeline for summarising documents too long for one pass
- Evaluate summaries with claim by claim checks rather than overlap scores alone
Lesson plan
- 1 Extractive and abstractive summaries Compare the two main approaches and the kinds of errors each one tends to make. Start
- 2 A field guide to summary failures Recognise the most common ways a summary misrepresents its source. Start
- 3 Purpose, audience and length Write summary requests that make the right things get kept and the right things dropped. Start
- 4 Summarising long documents Choose between chunk and combine, hierarchical and long context approaches. Start
- 5 Evaluating summaries honestly Measure faithfulness and coverage with methods that catch real errors. Start
- 6 Summaries in high stakes settings Decide when a summary needs human review and how to present its limits. Start
Try asking
About this tutor
For anyone who uses or builds automatic summaries and wants to trust them: analysts, product builders, researchers and curious users. You will learn the difference between extractive and abstractive summarisation, why a fluent summary can still be wrong, and how systems handle documents longer than a model can read at once. The lessons cover the typical failures (invented details, dropped caveats, wrong emphasis), ways to control length and audience, and how to evaluate summaries honestly, including the limits of overlap metrics and how to run faithfulness checks with people and with models. Examples range from meeting notes to research papers, with no coding required, though builders get pipeline guidance too.
Reviews
4.7
3 ratingsSample
- Ingrid H.Sample
I appreciated the honesty about ROUGE. Our team was optimising it and wondering why readers still complained.
- Ruth A.Sample
The claim labelling exercise changed how I review summaries at work. I found two unsupported numbers in a report summary the same week.
- Tariq M.Sample
Good mix of user and builder material. The map reduce lesson was what I needed for long contracts. I would have liked a worked pipeline example in more detail.
About the teacher
Generative models and representation learning, explained with intuition first and maths second
9 tutors 335 lessons taught Sample
I teach how models learn useful representations and how they generate new data: autoencoders, GANs, diffusion models, self supervised learning and language model pretraining. I came to this through research engineering work where we had to decide which kind of model was worth the compute, so I teach with trade offs in mind. Each topic starts with a picture or...
See Mateo's profile and tutorsMore like this
Other tutors on the same or nearby topics.