Unit Tests for Python AI Code
Test the deterministic parts of LLM apps properly, with fake model clients instead of flaky live calls.
A taste of a lesson
My tests call the real model and fail randomly. How do I fix that?
Split what you are testing. Your code (building the prompt, parsing the reply, retrying) should be tested with a fake client that returns replies you script, so results never vary. Pass the client into your function instead of creating it inside, then in tests give it a FakeClient whose method returns, say, a valid JSON string, a malformed one, or raises a rate limit error. Model quality is a separate question for an evaluation set. Which function in your code creates the client today?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Write readable tests with fixtures and parametrised cases
- Inject a fake model client to test prompts, parsing and retries
- Test failure paths such as malformed JSON, refusals and timeouts
- Separate unit tests of your code from evaluations of model quality
- Run fast tests on every change and keep live model tests apart
Lesson plan
- 1 Test basics in plain assert style Write, run and read simple tests and their failure output. Start
- 2 Fixtures and parametrised cases Reuse setup and cover many inputs without copying tests. Start
- 3 Making AI code testable Restructure code so the model client can be swapped for a fake. Start
- 4 Testing prompts and parsing Check that prompts are built correctly and outputs are parsed safely. Start
- 5 Testing failure paths Prove that retries, timeouts and refusals are handled as designed. Start
- 6 Where unit tests stop and evals begin Draw a clear line between code correctness and model quality. Start
Try asking
About this tutor
For developers building Python apps around model APIs who have few or no tests, or tests that call a live model and fail at random. You learn test basics in the plain assert style (discovery, fixtures, parametrised cases), then how to structure AI code so the parts that should be exact can be tested exactly: prompt building, output parsing, validation, retries, routing and error handling. You build a fake model client that returns chosen responses, including malformed JSON, refusals and timeouts. The tutor is clear about the line between unit tests, which check your code, and evaluations, which measure model quality.
Reviews
4.3
3 ratingsSample
- Marco V.Sample
Good material. I wanted a bit more on async tests, but the core approach transferred fine.
- Leon F.Sample
Injecting the client and writing a tiny fake class made my test suite fast and stable. Tests for the retry logic found a bug where we retried 400 errors forever.
- Zainab O.Sample
Clear line between unit tests and evals, which our team kept mixing up. Parametrised parsing tests were very practical.
About the teacher
Numerical Python and code quality for data and AI projects that have outgrown a single notebook
9 tutors 374 lessons taught Sample
I work on the part of AI projects nobody photographs: the arrays, dataframes, tests and packaging that let a prototype survive contact with real data. I came to this through scientific computing and later backend work, so I care about two things at once, getting the numbers right and keeping the code readable for the next person. My lessons are...
See Felix's profile and tutorsMore like this
Other tutors on the same or nearby topics.