Evaluating Agents: Tasks, Trajectories and Scores
Build an evaluation suite that tells you whether an agent change made things better or worse
A taste of a lesson
I changed my agent's prompt and the eval score went from 72% to 78%. Can I ship it?
Not yet. First, how many tasks, and how many runs per task? With 50 tasks run once, a six point jump could be noise. Rerun both versions three times on the same tasks and compare pass rates. Then look per task: did the new prompt fix eight tasks and break two? Those two breaks might matter more than the average, especially if they involve unsafe actions. Finally, check cost per task. Exercise: list the tasks that changed result between versions and sort them into fixed and broken.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Build a task suite from real usage that includes hard and adversarial cases
- Write programmatic outcome checks and validate model graded rubrics
- Grade trajectories for wasted steps and unsafe actions as well as outcomes
- Run tasks repeatedly and report pass rates and consistency honestly
- Compare agent versions fairly and keep the suite current
Lesson plan
- 1 What to measure in an agent Separate outcome, trajectory, cost and edge case behaviour as distinct things to evaluate. Start
- 2 Building a task suite Collect tasks that reflect real usage, including the awkward ones. Start
- 3 Checks a program can run Write outcome checks that do not depend on reading transcripts. Start
- 4 Model graded rubrics without self deception Use model graders for open ended outputs while knowing their limits. Start
- 5 Variance and fair comparisons Account for non deterministic runs when comparing agent versions. Start
- 6 Keeping the suite alive Maintain evaluation as the agent, users and models change. Start
Try asking
About this tutor
For engineers and applied researchers who need evidence, not impressions, about agent quality. Agents are harder to evaluate than single prompts because paths vary, runs are non deterministic and success can be partial. You will learn to build a task suite from real usage, define outcome checks that a program can run, grade trajectories for efficiency and safety, use model graded rubrics carefully, run each task several times to measure consistency, and compare versions fairly. By the end you will have a plan for an evaluation suite you can run before every prompt, tool or model change.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
I teach people to make AI agents safe, observable and testable before they are trusted with real work
9 tutors 290 lessons taught Sample
My background is in quality assurance and security testing, and I came to agents through the question every tester asks: how do we know this thing works, and what happens when it does not? I teach guardrails, permissions, approval steps, evaluation and trace debugging. My lessons are concrete. We look at a failure, name it, and design the check that...
See Ingrid's profile and tutorsMore like this
Other tutors on the same or nearby topics.