Jailbreaks, prompt injection and model misuse
Understand how AI systems are manipulated and the defensive design that limits the damage
A taste of a lesson
Why can't the AI just ignore instructions that come from a web page?
Because to the model, everything is text in one stream. The developer's instructions, your request and the web page content all arrive as tokens, and the model has learned to follow instructions wherever they appear. Training helps it prioritise the developer and user, but attackers find phrasings that slip through. So the robust fix is in system design: limit what the assistant can do after reading untrusted content, and require your confirmation before sensitive actions. Exercise: which actions of a browsing assistant would you put behind confirmation?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Distinguish jailbreaks, direct prompt injection and indirect prompt injection
- Explain why prompt injection cannot be fully solved at the model level today
- Threat model an AI system that has data access, tools and untrusted inputs
- Apply defence in depth: least privilege, confirmations, channel restrictions and monitoring
Lesson plan
- 1 Jailbreaks explained Understand what jailbreaks are and why they persist. Start
- 2 Prompt injection Understand direct and indirect injection. Start
- 3 When tools raise the stakes See why agents with data and actions are higher risk. Start
- 4 Defence in depth Design layered defences that assume some attacks succeed. Start
- 5 Testing and response Red team responsibly and handle incidents. Start
Try asking
About this tutor
For product people, developers, security staff and informed users who need to understand attacks on AI systems from a defensive point of view. You learn the difference between jailbreaks (getting a model to ignore its safety training) and prompt injection (hidden instructions in content that hijack an AI system), why both are hard to eliminate, and why the risk grows when models can browse, read email or take actions. You study defence in depth: least privilege, separating trusted and untrusted content, confirmation for consequential actions, output filtering, monitoring and red teaming. No attack recipes are taught; the focus is recognising and limiting risk.
Reviews
4.5
2 ratingsSample
- Nkechi U.Sample
Clear explanation of why injection is hard. The threat modelling exercise was the best part. More on monitoring would help.
- Bjorn A.Sample
The private data plus untrusted content plus exfiltration framing changed how we designed our internal assistant. Defensive focus throughout, no gimmicks.
About the teacher
I teach the big debates in AI, from alignment and safety testing to copyright, fairly and clearly
9 tutors 312 lessons taught Sample
I teach the questions people argue about: whether advanced AI could be dangerous, how labs test models, why models flatter us, who owns AI output, and how power over AI is distributed. I present each position in its strongest form before we examine it, and I separate evidence from speculation out loud. My background is in research communication and running...
See Bao's profile and tutorsMore like this
Other tutors on the same or nearby topics.