Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Reading AI benchmarks and research claims

Judge leaderboards, benchmark scores and model comparisons with a researcher's critical eye

By Bruno Azevedo Advanced How language models work 4.5(2) 19 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $9 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Reading AI benchmarks and research claims AI tutor following Bruno Azevedo's plan
Student:

Model A scores 87.2 and model B 86.5 on a benchmark. Is A better?

Tutor:

Not demonstrably. First, uncertainty: with a test set of, say, 1,000 items, the 95% interval at 87% is about plus or minus 2 points, so 0.7 is well within noise. Second, settings: were both run with the same prompts, examples and answer extraction, by the same party? Third, contamination and relevance: does the benchmark reflect your task at all? A 0.7 gap rarely justifies a decision. Exercise: compute the rough interval for 300 items at 87% using sqrt(p(1-p)/n).

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain what benchmarks measure and the main ways scores mislead
  • Identify contamination, saturation, prompt sensitivity and selective reporting
  • Interpret preference leaderboards and model graded evaluations with their biases
  • Read a technical report for baselines, variance and evaluation settings
  • Design a small task specific evaluation to complement public benchmarks

Lesson plan

6 lessons. Pick one to start there.

  1. 1 What a benchmark is and is not Understand benchmarks as proxies with specific scopes. Start
  2. 2 Contamination and saturation Recognise when scores are inflated or no longer informative. Start
  3. 3 Settings, prompts and variance See how evaluation choices and sample size change results. Start
  4. 4 Preference leaderboards and model judges Interpret crowd votes and model graded scores critically. Start
  5. 5 Reading the report Extract the details that determine how much a result means. Start
  6. 6 Build your own evaluation Design a small, task specific evaluation that informs real decisions. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For technically literate readers, analysts, journalists, investors and engineers who need to evaluate claims like 'new model tops leaderboard'. You learn what benchmarks measure and what they miss, how test set contamination, saturation, prompt sensitivity and selective reporting distort scores, how human preference leaderboards work and their biases, and how to read a technical report or paper for the details that matter: baselines, variance, ablations and evaluation settings. You practise on realistic invented announcements and finish with a checklist for deciding how much weight a result deserves and how to run a small task specific evaluation of your own.

Reviews

4.5

2 ratingsSample

  • Anders V.Sample

    The confidence interval exercise was humbling: half the 'wins' in our vendor comparison were noise. Excellent on preference leaderboard biases.

  • Shreya K.Sample

    Dense but well structured. The build your own eval lesson was the most practical. More worked examples on contamination detection would help.

About the teacher

Bruno Azevedo

I teach people to judge AI claims, spot synthetic media and report on AI without the hype

9 tutors 4.5(21) 435 lessons taught Sample

I teach media literacy for the age of AI. My learners include journalists, students, sceptics and anyone tired of breathless headlines in both directions. We practise reading claims about AI critically, checking images and video, understanding why AI text detectors fail, and asking the questions a careful reporter would ask. My background is in newsroom fact checking and training reporters,...

See Bruno's profile and tutors