Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.

Find a tutor

657 tutors in 31 topics, built by 75 teachers. Each one follows a lesson plan its teacher wrote.

Filters

Clear
Evaluating a model after fine tuning

Evaluating a model after fine tuning

Prove a tuned model is actually better, on your task and everywhere else it mattersIntermediateEvaluation and testing4.7(3)65 lessonsSample
Neha Varadan$8
Logging and Tracing LLM Calls

Logging and Tracing LLM Calls

Record every model call and pipeline step so you can explain cost, slowness and bad answers.IntermediateBuilding with LLM APIs4.0(3)65 lessonsSample
Farid Haddad$8
Unit Tests for Python AI Code

Unit Tests for Python AI Code

Test the deterministic parts of LLM apps properly, with fake model clients instead of flaky live calls.IntermediateEvaluation and testing4.3(3)61 lessonsSample
Felix Brandt$6
Cost and Latency Budgets for LLM Features

Cost and Latency Budgets for LLM Features

Set cost and speed targets for an LLM feature, measure them honestly and trade them against quality.All levelsBuilding with LLM APIs4.5(2)52 lessonsSample
Farid Haddad$8
Rubric Grading for Open Ended Output

Rubric Grading for Open Ended Output

Design rubrics that make grading emails, summaries and explanations consistent, fair and repeatable.IntermediateEvaluation and testing4.3(3)49 lessonsSample
Gonzalo Ibarra$6
Measuring Hallucination in LLM Apps

Measuring Hallucination in LLM Apps

Measure how often your app states unsupported things, with claim level checks you can repeat and trust.IntermediateEvaluation and testing4.5(2)42 lessonsSample
Gonzalo Ibarra$8
Regression Testing Your Prompts

Regression Testing Your Prompts

Catch quality drops before users do by running an eval suite on every prompt, model or setting change.IntermediateEvaluation and testing4.5(2)42 lessonsSample
Gonzalo Ibarra$7
CI and testing for ML projects

CI and testing for ML projects

Add fast automated checks that catch broken data, code and models before they reach usersIntermediateEvaluation and testingNew
Malik Brennan$8
Human Review That Scales

Human Review That Scales

Set up human review of AI outputs that is consistent, affordable and feeds back into better systems.All levelsEvaluation and testingNew
Gonzalo Ibarra$5
Evaluating generated text

Evaluating generated text

Measure the quality of generated text with metrics, people and model judges, and know each one's limitsIntermediateEvaluation and testingNew
Mateo Rojas$8