Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Evaluating Agents: Tasks, Trajectories and Scores

Build an evaluation suite that tells you whether an agent change made things better or worse

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Evaluating Agents: Tasks, Trajectories and Scores AI tutor following Ingrid Solberg's plan
Student:

I changed my agent's prompt and the eval score went from 72% to 78%. Can I ship it?

Tutor:

Not yet. First, how many tasks, and how many runs per task? With 50 tasks run once, a six point jump could be noise. Rerun both versions three times on the same tasks and compare pass rates. Then look per task: did the new prompt fix eight tasks and break two? Those two breaks might matter more than the average, especially if they involve unsafe actions. Finally, check cost per task. Exercise: list the tasks that changed result between versions and sort them into fixed and broken.

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Build a task suite from real usage that includes hard and adversarial cases
  • Write programmatic outcome checks and validate model graded rubrics
  • Grade trajectories for wasted steps and unsafe actions as well as outcomes
  • Run tasks repeatedly and report pass rates and consistency honestly
  • Compare agent versions fairly and keep the suite current

Lesson plan

6 lessons. Pick one to start there.

  1. 1 What to measure in an agent Separate outcome, trajectory, cost and edge case behaviour as distinct things to evaluate. Start
  2. 2 Building a task suite Collect tasks that reflect real usage, including the awkward ones. Start
  3. 3 Checks a program can run Write outcome checks that do not depend on reading transcripts. Start
  4. 4 Model graded rubrics without self deception Use model graders for open ended outputs while knowing their limits. Start
  5. 5 Variance and fair comparisons Account for non deterministic runs when comparing agent versions. Start
  6. 6 Keeping the suite alive Maintain evaluation as the agent, users and models change. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers and applied researchers who need evidence, not impressions, about agent quality. Agents are harder to evaluate than single prompts because paths vary, runs are non deterministic and success can be partial. You will learn to build a task suite from real usage, define outcome checks that a program can run, grade trajectories for efficiency and safety, use model graded rubrics carefully, run each task several times to measure consistency, and compare versions fairly. By the end you will have a plan for an evaluation suite you can run before every prompt, tool or model change.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Ingrid Solberg

I teach people to make AI agents safe, observable and testable before they are trusted with real work

9 tutors 4.5(17) 290 lessons taught Sample

My background is in quality assurance and security testing, and I came to agents through the question every tester asks: how do we know this thing works, and what happens when it does not? I teach guardrails, permissions, approval steps, evaluation and trace debugging. My lessons are concrete. We look at a failure, name it, and design the check that...

See Ingrid's profile and tutors