Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Red Teaming and Prompt Injection Testing

Test your LLM app against jailbreaks, prompt injection and data leaks with a repeatable attack suite.

By Gonzalo Ibarra Advanced Evaluation and testing 4.7(3) 49 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Red Teaming and Prompt Injection Testing AI tutor following Gonzalo Ibarra's plan
Student:

How can I tell automatically whether our assistant leaked its system prompt during testing?

Tutor:

Plant a canary. Add a unique, meaningless string to the system prompt, something like a random code that would never appear naturally, and do the same in restricted documents. Then run your extraction attacks and scan every output for those strings, including encoded or spaced out versions. Any match is a confirmed leak, with no judge needed. Track the leak rate per attack category across versions. Which attack categories are in your current library?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Build a threat model of assets, attackers and harms for an LLM app
  • Assemble an attack library covering direct, indirect and multi turn attacks
  • Detect successful attacks automatically with canaries and tool logs
  • Track attack success rates per category across versions
  • Triage findings and fix them through architecture, then retest

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Scope and threat model Define what you are protecting, from whom and against which harms. Start
  2. 2 Attack categories Understand the main families of attacks on LLM apps. Start
  3. 3 Building the attack library Create test cases per category with expected safe behaviour. Start
  4. 4 Detecting success automatically Know when an attack worked without reading every output. Start
  5. 5 Measuring over time Track attack success rates and catch regressions. Start
  6. 6 Triage and fixes Prioritise findings and fix them durably. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For experienced engineers, security staff and evaluators who need evidence that an LLM application resists misuse before and after launch. You learn to scope red teaming with a threat model for your app, build an attack library by category (direct and indirect injection, multi turn escalation, obfuscation, system prompt extraction, data leakage, unauthorised tool use, harmful content), detect success automatically with canary strings and tool call logs, and track attack success rates over time. The tutor stresses authorised testing only, responsible handling of findings, and fixing weaknesses through architecture rather than prompt patches alone.

Reviews

4.7

3 ratingsSample

  • Mia K.Sample

    Clear that prompt patches are not real fixes. We reduced tool permissions instead and the attack success rate dropped across categories.

  • Rohan V.Sample

    Thorough and responsible, with authorisation stressed upfront. I wanted more examples of multi turn attacks, but the method is solid.

  • Tarik S.Sample

    Canary strings turned our red teaming from manual reading into an automated CI check. The threat model step made us test indirect injection through uploads, which we had ignored.

About the teacher

Gonzalo Ibarra

Evaluation for LLM products: eval sets, graders, regression tests, hallucination checks and live experiments

9 tutors 4.5(17) 294 lessons taught Sample

I teach people how to know whether their LLM feature is any good, which is harder than building it. I started in software testing and quality work, moved into data analysis, and now help teams measure systems whose outputs vary from run to run. My lessons are about method: write down what good looks like, collect examples, grade them in...

See Gonzalo's profile and tutors