Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Prompt Injection and Agent Security

Understand how untrusted text hijacks agents and design systems that limit the damage when it does

By Ingrid Solberg Advanced AI agents 4.7(3) 45 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $13 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Prompt Injection and Agent Security AI tutor following Ingrid Solberg's plan
Student:

Our agent summarises customer emails and can also send replies. Could an email trick it into doing something bad?

Tutor:

Yes, assume it can. An email is untrusted content, and hidden instructions in it might ask the agent to send data elsewhere or reply with something harmful. Since it can send replies, you have an outbound channel. Reduce the damage: let replies go only to the original sender, require approval before anything is sent, and make sure the agent cannot look up other customers' data while handling one email. Then test with planted instructions in sample emails. What data can the agent reach while it is processing a single email today?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain direct and indirect prompt injection and why models are vulnerable
  • Identify the combination of private data, untrusted content and outbound actions
  • Design architectural controls that limit damage when injection succeeds
  • Build a repeatable set of injection test cases for your own agent
  • Threat model an agent and prioritise fixes by impact

Lesson plan

6 lessons. Pick one to start there.

  1. 1 What prompt injection is and why it works Understand direct and indirect injection and why models cannot reliably separate data from instructions. Start
  2. 2 The dangerous combination Recognise when private data, untrusted input and outbound actions together create real risk. Start
  3. 3 Architecture over wording Design system structure so a hijacked model cannot do serious harm. Start
  4. 4 Controls on actions and outputs Limit what an agent can send or change even when its instructions are subverted. Start
  5. 5 Testing for injection Build and maintain a test set of injection attempts in the content your agent reads. Start
  6. 6 Threat modelling your agent Map data, inputs and actions for your agent and choose the fixes that matter most. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers and security minded builders whose agents read content they do not control: web pages, emails, documents, tickets, tool outputs or files from users. Prompt injection is the defining security problem of agents: instructions hidden in data can steer the model into leaking information or taking actions nobody asked for. There is no complete fix today, so this tutor teaches defence in depth. You will study direct and indirect injection, data exfiltration paths, the dangerous combination of private data, untrusted content and outbound actions, and the architectural controls that reduce risk. You will finish by threat modelling your own agent.

Reviews

4.7

3 ratingsSample

  • Ravi T.Sample

    Our injection test set started from the exercise in lesson five. It caught a regression after a model change.

  • Julia M.Sample

    Honest about the limits, which I appreciated after reading too many 'just add this prompt' posts. The threat modelling lesson was dense but worth it.

  • Sami E.Sample

    The private data, untrusted content, outbound action triangle reorganised our whole review. We removed auto rendered links the next day.

About the teacher

Ingrid Solberg

I teach people to make AI agents safe, observable and testable before they are trusted with real work

9 tutors 4.5(17) 290 lessons taught Sample

My background is in quality assurance and security testing, and I came to agents through the question every tester asks: how do we know this thing works, and what happens when it does not? I teach guardrails, permissions, approval steps, evaluation and trace debugging. My lessons are concrete. We look at a failure, name it, and design the check that...

See Ingrid's profile and tutors