Classification Metrics Beyond Accuracy
Read a confusion matrix and pick the metric that matches the cost of each mistake
A taste of a lesson
Our fraud model has 99.4 percent accuracy. My manager is thrilled. Should I be?
Not yet. First ask the fraud rate. If 0.6 percent of transactions are fraud, a model that never flags anything also gets 99.4 percent accuracy. You need the confusion matrix: how many frauds did it catch (recall) and how many of its alerts were real (precision)? Then ask at what threshold, and how many alerts per day the team can review. Try this: if there are 60 frauds in 10,000 transactions and the model catches 30 with 90 false alarms, what are precision and recall?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Build a confusion matrix from counts and read it fluently
- Calculate precision, recall, specificity and F1 by hand
- Explain why accuracy misleads on imbalanced data
- Choose between ROC and precision recall curves for a problem
- Pick a threshold from error costs and team capacity
Lesson plan
- 1 The confusion matrix Lay out every prediction as a true or false positive or negative. Start
- 2 Why accuracy misleads See how imbalanced classes make accuracy look good for bad models. Start
- 3 Precision, recall and F1 Calculate and interpret the three most used classifier metrics. Start
- 4 Curves across thresholds Use ROC and precision recall curves to compare models. Start
- 5 Choosing a threshold Set the cut off from costs, capacity and goals. Start
- 6 Calibration and multiclass Check probability quality and summarise multiclass results. Start
Try asking
About this tutor
A beginner tutor for anyone who has to judge a classifier, whether you build models or receive reports about them. You will build confusion matrices by hand, calculate precision, recall, F1 and specificity, and see why accuracy can look excellent while a model is useless. Later lessons explain ROC curves and AUC, precision recall curves for rare events, threshold choice and calibration, all with small realistic examples like spam filters and fraud alerts. You finish able to ask the right questions when someone shows you a single impressive number.
Reviews
4.7
3 ratingsSample
- Samuel T.Sample
Clear and practical. The precision recall versus ROC lesson was the most useful. Calibration part was a bit short.
- Beatriz C.Sample
I review vendor models at work and now I always ask for the base rate and the threshold. The spam example with the always-negative baseline was eye opening.
- Mei L.Sample
Calculating everything from counts first made the formulas stick. Good for beginners.
About the teacher
Model evaluation you can trust: splits, metrics, leakage and tuning
9 tutors 439 lessons taught Sample
Most of the machine learning failures I have seen were not about the algorithm. They came from a test set that was used too often, a feature that leaked the answer, or a metric that hid the real cost of mistakes. I teach the unglamorous discipline that makes model results believable: how to split data, how to validate, which metric...
See Lukas's profile and tutorsMore like this
Other tutors on the same or nearby topics.