Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Moderation, Refusals and Safe Responses

Design how your LLM app handles harmful requests, refusals and sensitive topics, without blocking normal users.

By Greta Lindqvist All levels Building with LLM APIs 4.3(3) 49 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $5 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Moderation, Refusals and Safe Responses AI tutor following Greta Lindqvist's plan
Student:

Our app for nurses keeps refusing questions about medication doses. How do we fix it?

Tutor:

Start by giving the model the context it lacks. In the system prompt, state that users are verified nursing staff using the tool for professional reference, that dosing questions are expected, and how answers should point to local protocols and pharmacist checks. That often reduces needless refusals within provider policies. Then measure: collect thirty real, legitimate dosing questions and see how many are refused before and after. Also check whether your own moderation layer, not the model, is blocking them. Where does the refusal text come from?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Describe the layers of safety in an LLM application
  • Write a short content policy suited to your users
  • Detect refusals and present them helpfully
  • Measure over blocking and under blocking on your own data
  • Handle sensitive situations with care and set up review of flagged cases

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Layers of safety Map model behaviour, classifiers, app rules and human review in your system. Start
  2. 2 Writing your policy Draft a one page policy for your users and use case. Start
  3. 3 Refusals users can live with Detect refusals and turn them into helpful responses. Start
  4. 4 Sensitive situations Respond with care when users show signs of crisis or vulnerability. Start
  5. 5 Measuring your filters Test safety layers on realistic inputs from your own domain. Start
  6. 6 Review and improvement Run a light review loop that improves policy and prompts over time. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For developers, product owners and trust and safety staff responsible for an LLM feature used by the public or by staff. You learn the layers of safety in an app: the model's own behaviour, moderation classifiers on input and output, and your own written policy. You practise writing a short content policy for your audience, detecting and presenting refusals well, handling sensitive situations such as signs of self harm with care, measuring over blocking and under blocking on your own data, and setting up review of flagged cases. The tutor is clear that no filter is perfect and that policy decisions belong to the organisation.

Reviews

4.3

3 ratingsSample

  • Hannah B.Sample

    Writing a one page policy first changed our conversations from vague worry to specific decisions. The crisis response lesson was handled thoughtfully.

  • Laura C.Sample

    Useful for a non engineer like me on our trust and safety team. Clear that legal decisions sit with us, not the tutor.

  • Emre Y.Sample

    We measured false blocks on our own data and found our keyword list caused most of them, not the model. Very practical.

About the teacher

Greta Lindqvist

Structured output, tool calling and safe input handling for LLM applications that must behave predictably

9 tutors 4.5(18) 308 lessons taught Sample

I teach the parts of LLM apps where free text has to meet real software: JSON that must parse, tools the model calls, images and documents coming in, and users who send things you did not plan for. I spent years writing integrations between messy systems, which taught me to treat every input as untrusted and every output as something...

See Greta's profile and tutors