Classifying Text with an LLM API
Sort tickets, reviews or messages into categories with a model, and measure how accurate it really is.
A taste of a lesson
My ticket classifier is 92 percent accurate. Is that good enough to launch?
It depends on where the 8 percent of errors land. If 'urgent outage' tickets are 3 percent of traffic, the classifier could miss every one of them and still score 92 percent. Look at per category recall, especially for categories where a miss is costly, and at the confusion matrix to see what gets mixed up. Also check the test set reflects real traffic. What recall do you get for your most important category?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Write clear category definitions with boundary rules and an other option
- Build a labelled test set and check human agreement
- Write a classification prompt with constrained, validated labels
- Measure accuracy, per category precision and recall, and confusion
- Decide when to route items to people or switch to a cheaper classifier
Lesson plan
- 1 Defining categories Write definitions and boundary rules that people and models can apply. Start
- 2 A labelled test set Collect and label real examples to measure against. Start
- 3 The classification prompt Write a prompt that produces only valid labels. Start
- 4 Measuring results Compute accuracy, precision, recall and a confusion matrix. Start
- 5 Improving and routing Fix the biggest confusions and send uncertain items to people. Start
- 6 Scale and maintenance Keep cost under control and catch drift over time. Start
Try asking
About this tutor
For beginners who want to use a model API to label text: routing support tickets, tagging feedback, detecting topics or urgency. You learn to define categories clearly with boundary rules, write a classification prompt with examples, constrain the answer to valid labels, and, most importantly, measure accuracy on a labelled test set with per category results and a confusion matrix. The tutor covers single versus multi label tasks, an honest other category, why self reported confidence is unreliable, and when a smaller model or a traditional classifier is the better choice at high volume.
Reviews
4.5
2 ratingsSample
- Mariana S.Sample
We were proud of our accuracy until the confusion matrix showed refund requests going to billing. Boundary rules fixed most of it. Great beginner explanation of precision and recall.
- Jakob T.Sample
Labelling a test set felt slow but it is the most valuable thing we did. Good point about model confidence numbers not meaning much.
About the teacher
Structured output, tool calling and safe input handling for LLM applications that must behave predictably
9 tutors 308 lessons taught Sample
I teach the parts of LLM apps where free text has to meet real software: JSON that must parse, tools the model calls, images and documents coming in, and users who send things you did not plan for. I spent years writing integrations between messy systems, which taught me to treat every input as untrusted and every output as something...
See Greta's profile and tutorsMore like this
Other tutors on the same or nearby topics.