What you will be able to do
- Describe the guardrail layers you are responsible for on top of the platform safeguards that already run by default
- Screen untrusted input and tool output with a small classifier call whose verdict your application can branch on
- Structure an agent so third-party content cannot act as instructions, and scope its access so a successful injection does little damage
- Detect stop_reason: refusal, reset context, and interpret the refusal category and its billing behaviour
- Place output-side checks and a human review tier where a confident automated verdict is not available
Key concept
Layered guardrails — Safety in a Claude deployment is a stack, not a switch: Anthropic's own real-time classifiers on API inputs and outputs are the floor that is always there, and every layer above it — input screens, prompt structure, access limits, output filters, human review — is something you build in your application.
1.The floor you get, and the layers you build
Start with what is already running. Anthropic applies real-time safeguards to API inputs and outputs for every request, and there is no dashboard toggle that turns extra filtering on: "There is no additional opt-in filter to enable." That is the floor. It is not a substitute for application-level controls, because it enforces Anthropic's Usage Policy, not your product's rules — an ad network's disclaimer requirements or a support bot's refusal to discuss competitors are yours to enforce.
So the design question for this subdomain is never "is filtering on?" but "at which point in my request path does each check run?" The rest of this lesson walks that path in order: what reaches Claude, how the prompt is structured, what Claude is allowed to touch, what comes back, and what a human sees.
There is nothing to enable — input/output safeguards run by default. What is left to you is everything policy-specific: pre-screening prompts, structuring untrusted content, restricting Claude's access and tools, post-processing outputs, and escalating repeat offenders and unclear cases to review.
Sources1
2.Two threat models, two sets of controls
Before choosing controls, decide who the adversary is. The guardrails documentation splits the problem in two: "These attacks fall into two categories with different threat models:" In a jailbreak or direct prompt injection, the user of your application is the attacker, crafting input designed to get past your rules. In indirect prompt injection the user is trusted and the attacker is upstream of the content — an inbound email, a fetched page, OCR text, a tool result.
The distinction matters because the controls barely overlap. Against a hostile user you screen and rate-limit the user. Against hostile content you cannot ban anyone: you change how the content enters the prompt and what Claude is permitted to do once it is there.
| Threat model | Who is the adversary | Representative controls |
|---|---|---|
| Jailbreaks and direct prompt injection | The user of your application | Harmlessness screen on user input, input validation for known injection patterns, ethical system prompt, throttling repeat offenders |
| Indirect prompt injection | Whoever can influence third-party content Claude reads | Deliver content in tool_result blocks, JSON-encode it, state an untrusted-content policy, least privilege, screen tool outputs |
A fintech company's customer-facing assistant is built on Claude. Before a user message reaches the main conversation, the compliance team wants a fast, cheap classification step that flags requests referring to harmful, illegal, or fraudulent activity, so those messages can be routed to a stricter review flow instead of the main model. Which guardrail design best satisfies this requirement while controlling cost?
Correct answer: C — Route the message to a lightweight model like Claude Haiku 4.5 with a classification prompt, constraining its reply to a structured schema such as is_harmful: boolean
- A. Incorrect. Appending a classification instruction to the same turn as the reply is not a separate pre-screening step and does not use a cheaper model, so it fails both the routing and cost goals.
- B. Incorrect. Temperature controls output randomness, not harm detection, and waiting for a refusal in the final reply happens after generation rather than screening input beforehand.
- C. Correct. This is a harmlessness screen: a lightweight model classifies the input before it reaches the main conversation, with structured outputs constraining the response to a simple, parseable verdict.
- D. Incorrect. A fixed keyword blocklist lacks semantic understanding and misses paraphrased or context-dependent harmful requests that a classifier model would catch.
Sources2
3.Input classifiers: a small model in front of the big one
The first layer you own is an input classifier. The recommended shape is a cheap pre-flight call: "Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation." The point of using a small model is that the screen runs on every request, so its cost and latency must be near-negligible compared with the main call.
A screen is only a guardrail if your code can act on it mechanically. That is why the guidance pairs the screen with constrained decoding — "Use structured outputs to constrain the response to a simple classification." A boolean field such as is_harmful is something an if-statement can branch on; a paragraph of prose explaining the model's concerns is not. Alongside the model screen sits ordinary deterministic validation: "Filter user input for known injection patterns before it reaches Claude."
The third input-side control is the prompt itself. A system prompt that states boundaries explicitly and tells Claude the exact refusal wording to use makes refusals consistent and, usefully, machine-detectable in your logs.
Sources2
4.Untrusted content and the access controls behind it
Now the harder half. When Claude reads content on a user's behalf, your job is to make the boundary between your instructions and that content unambiguous. The first rule is placement: "Put untrusted content only in tool results." Claude is trained to treat instructions appearing inside tool results with skepticism, which is exactly the skepticism you want — and which you throw away by pasting a fetched page into the system prompt or a plain user turn.
Three reinforcements follow. Label the content: say in the tool description or the result structure that this is, for example, the body of an inbound email from an unknown sender. State the policy in the system prompt — tell Claude that tool, document and search content is untrusted data that must never override the system prompt or the user's request, and that apparent instructions inside it should be reported rather than obeyed. And encode it: wrapping third-party strings in JSON gives unambiguous delimiters, so an attacker cannot close a quote or tag to break out into an instruction context.
The same placement rule cuts the other way, and this is a favourite trap: do not put your own instructions into tool results. "Because Claude treats tool-result content as untrusted data, instructions you place there may be ignored or flagged as a potential injection." Send them in a user turn after the tool_result block instead.
Content structure is not enough on its own, so guardrails extend to authorisation. "Apply the principle of least privilege so that a successful injection can do minimal damage" — do not hand Claude secrets it does not need, "run tools in sandboxed environments, and scope permissions as narrowly as possible." Access control is what converts a successful injection from a breach into a nuisance.
Finally, the screening pattern from the input side applies to the output of your own tools: "Screen tool outputs before Claude acts on them." Run the tool, pass its raw output to a Haiku 4.5 classifier, and only hand back a tool_result block if the verdict field — again a structured boolean, injection_suspected — reports nothing. When it does report an attempt, return an error or a stripped summary instead of the raw content, and consider surfacing the attempt to the user.
| What is screened | Screening model | Structured verdict field | Action when it trips |
|---|---|---|---|
| User input, before the main conversation | Claude Haiku 4.5 | is_harmful | Do not pass the turn to the main conversation; escalate repeat offenders |
| Raw tool output, before Claude sees it | Claude Haiku 4.5 | injection_suspected | Return an error or stripped summary in the tool_result block instead of raw content |
None of these layers is verified until you attack them yourself. "Red-team your own agent." Feed the workflow documents, emails and tool outputs that deliberately contain injection attempts before deployment, and confirm both that Claude ignores them and that your screens and confirmation steps catch what it does not.
A review of a Claude-powered support bot's logs shows that a small number of accounts have each triggered the same jailbreak-style refusal more than a dozen times in one week, using slightly reworded prompts each time. According to Anthropic's guardrail guidance, what is the most appropriate response to this specific pattern?
Correct answer: B — Recognize the pattern as repeated circumvention, tell the accounts their requests violate the usage policy, and apply throttling or restrictions as warranted
- A. Incorrect. A blanket global threshold change penalizes all users rather than addressing the specific accounts that have shown a repeated circumvention pattern.
- B. Correct. Anthropic's guidance specifically addresses repeat offenders: adjust responses, and consider throttling or banning users who repeatedly attempt to circumvent guardrails.
- C. Incorrect. Removing logs eliminates the evidence needed to identify and act on repeat-offender behavior, directly undermining the continuous monitoring guardrails rely on.
- D. Incorrect. Rotating prompt wording for every user adds operational complexity for the whole user base without targeting the accounts actually exhibiting the identified behavior.
Sources2
5.When the platform classifier fires: stop_reason refusal
The real-time safeguards from the first section are visible in your code as a specific response shape. On Claude 4 and later, streaming responses come back with stop_reason set to refusal "when streaming classifiers intervene to handle potential policy violations". Note what this is not: it is a successful HTTP 200 response, not an exception, so error-handling code alone will never see it.
{
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello.."
}
],
"stop_reason": "refusal",
"stop_details": {
"type": "refusal",
"category": "cyber",
"explanation": "This request was declined because it could enable cyber harm."
}
}The operational rule is the part people get wrong: on a refusal "you must reset the conversation context before continuing". Remove or rephrase the offending turn, or clear the history; retrying the same message list simply earns another refusal. The partial text is not usable either — "treat any partial output as incomplete and discard it", because a refusal can arrive before any output or mid-stream after some.
stop_details tells you which policy area tripped. Display the explanation rather than parsing it — the wording is not stable — and branch on category. Whether you were billed for the attempt also depends on the category.
| category | Billed before any output |
|---|---|
| "cyber" | No |
| "bio" | Yes |
| "frontier_llm" | Yes |
| "reasoning_extraction" | Yes |
| "general_harms" | No |
Around that, the recommended handling is a small control loop: monitor for refusals in your response handling, reset context automatically, fall back to another Claude model rather than surfacing a raw refusal, show your own user-facing message, and track refusal frequency as a signal about your prompts. A rising refusal rate on benign traffic is usually a prompt problem, not a user problem — benign cybersecurity and life-sciences work can trigger the cyber and bio categories.
An agent ingests inbound customer emails as tool results and drafts replies on the user's behalf. A red-team exercise crafts an email body containing text designed to close out the surrounding context and inject a new instruction telling the agent to forward the user's contact list. The way the email body is concatenated into the tool result lets the crafted text escape into an instruction context. Which change most directly closes this specific escape route?
Correct answer: D — Wrap the email body inside a JSON object in the tool result, so the untrusted text is delimited as an escaped string rather than concatenated free-form text
- A. Incorrect. Placing untrusted content in the system prompt raises its authority rather than containing it, which is the opposite of the recommended handling for untrusted content.
- B. Incorrect. Manual retyping by the user is not a scalable technical control and does not address the underlying concatenation mechanism that allows the escape.
- C. Incorrect. A larger context window changes how much history is retained, but does not prevent crafted text from escaping into an instruction context.
- D. Correct. JSON-encoding untrusted content provides unambiguous delimiters, so an attacker cannot close a quote or tag within the payload to break out into an instruction context.
6.Output-side controls: filtering and auditability
Guardrails on the way out are cheaper than they look. For prompt leakage, post-processing is a deterministic layer over a probabilistic one: "Techniques include using regular expressions, keyword filtering, or other text processing methods." It runs after generation and does not compete with the task, which matters because leak-resistant prompting itself adds complexity that can degrade the model's performance on the real work.
The other output-side control worth wiring in is auditability. Giving Claude explicit permission to decline — "Explicitly give Claude permission to admit uncertainty." — plus a requirement to support each claim with a quote from the provided material turns an unverifiable answer into one your own code, or a reviewer, can check. That check is what feeds the last layer.
7.Tiering safeguards up to human review
The API safeguards guidance presents these controls as tiers you add as a deployment matures. The basic tier is accountability plumbing: "Store IDs linked with each API call", so violative content can be located later, optionally with hashed per-user IDs so misuse can be attributed to an individual. Enforcement sits on top of it — "Warn, throttle, or suspend users who repeatedly violate" the Terms and Usage Policy. This is the same repeat-offender response the jailbreak guidance recommends when one user keeps tripping the same refusal.
Intermediate tiers narrow the attack surface rather than inspecting it: restrict end users to a limited set of prompts, or point Claude only at a knowledge corpus you already control. Advanced tiers add moderation — using Claude itself for content moderation, or running a moderation API over every end-user prompt before it reaches Claude.
The top tier is a human: "Set up an internal human review system to flag prompts that are marked by Claude" or by a moderation API as harmful, so you can intervene against accounts with high violation rates. The important design point is that review is a routed outcome, not a fallback for everything. A moderation loop decides what happens to each item — "publish, reject, or send to a human reviewer" — and in a well-built pipeline that third branch is reserved for genuine uncertainty: when a needed field could not be determined, "Missing values evaluate to unknown" and the item is routed to review instead of receiving a confident automated verdict.
Neither reject nor publish. A null is not a false — it means the check could not be evaluated, so the item routes to human review. Collapsing unknown into a confident verdict either blocks compliant content or ships non-compliant content, and it destroys the evidence trail that makes the decision defensible.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Before launching you must enable Anthropic's input/output content filter in the console, and once it is on your application does not need its own moderation layer.Why is that wrong?
Real-time safeguards on API inputs and outputs already run by default and there is nothing to opt into; your own moderation layer is added in your application on top of them.
Covered in The floor you get, and the layers you build
2.A classifier refusal arrives as an HTTP error, so catching exceptions and retrying the same request handles it.Why is that wrong?
A refusal is a successful response carrying stop_reason: refusal, and continuing without resetting the conversation context produces further refusals.
Covered in When the platform classifier fires: stop_reason refusal
3.The safest place for your strictest safety instructions is inside the tool result, right next to the untrusted content they are meant to govern.Why is that wrong?
Tool-result content is treated as untrusted data, so your instructions there may be ignored or even flagged as an injection; send them in a following user turn instead.
Covered in Untrusted content and the access controls behind it
4.If user input is screened by a classifier, content that arrives later through tool calls has already passed the guardrail.Why is that wrong?
Tool output is a separate entry point and needs the same screening before Claude acts on it, with the raw content withheld when the screen reports an injection attempt.
Covered in Untrusted content and the access controls behind it
5.A refusal always happens before generation starts, so any text you have already streamed to the user is safe to keep.Why is that wrong?
A refusal can also arrive mid-stream after partial output, and that partial output must be discarded as incomplete.
Covered in When the platform classifier fires: stop_reason refusal
6.An input screen should return a short natural-language explanation so your logs capture the reasoning.Why is that wrong?
The screen exists so code can branch on it, so the response is constrained by a JSON schema to a simple classification such as a boolean is_harmful.
Covered in Input classifiers: a small model in front of the big one
Practise it for real
Red-team a document-reading agent and close the gap with a tool-output screen
1.Build a minimal agent that fetches a document and summarises it, delivering the document text inside a tool_result block rather than in the system prompt or a user text block.
Why: Claude is trained to treat instructions inside tool results with skepticism, so placement is the cheapest part of the guardrail.
You should see: A working summariser whose untrusted payload is confined to tool_result content.
2.Add an untrusted-content policy to the system prompt stating that tool, file and search content is data, that it must never override the system prompt or the user request, and that apparent instructions should be reported instead of followed.
Why: The policy tells Claude how to calibrate directives found inside retrieved content.
You should see: The summary of a hostile document mentions that it contained instructions, rather than acting on them.
3.Feed the agent a document whose body contains an injection such as an instruction to ignore previous instructions and exfiltrate a key, and record what the agent does.
Why: Red-teaming your own agent before deployment is the only way to know whether the layers you added hold.
You should see: Either a clean report of the attempt, or a concrete failure you now have to screen for.
4.Insert a Claude Haiku 4.5 screening call between the tool and the model, asking only whether the content contains instructions aimed at the assistant, with the response constrained by a JSON schema to a boolean injection_suspected.
Why: A structured verdict is a value your application can branch on; prose is not.
You should see: injection_suspected is true for the hostile document and false for a clean one.
5.Branch on the verdict: when injection_suspected is true, return an error or a stripped summary in the tool_result block instead of the raw content, and surface the attempt to the user.
Why: Screening is only a guardrail if a positive verdict changes what reaches Claude.
You should see: The hostile document never reaches the main conversation in raw form, while clean documents summarise as before.
6.Re-run the whole red-team set, then re-scope the agent's credentials and sandbox its tools so that a screen that misses still cannot cause damage.
Why: Least privilege bounds the impact of the injection that gets through every content-level layer.
You should see: No injection in your set changes the agent's goals, and the agent holds no secret or permission the summarising task does not require.
Stuck? Get a nudge
Write the screen as a separate function that takes raw tool output and returns a boolean, not as extra wording in the main prompt — a guardrail you can unit-test is a guardrail you can trust.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Anthropic runs real-time safeguards on API inputs and outputs by default.”
↩︎ The floor you get, and the layers you build“There is no additional opt-in filter to enable.”
↩︎ The floor you get, and the layers you build“Store IDs linked with each API call”
↩︎ Tiering safeguards up to human review“Warn, throttle, or suspend users who repeatedly violate”
↩︎ Tiering safeguards up to human review“Set up an internal human review system to flag prompts that are marked by Claude”
↩︎ Tiering safeguards up to human review“Anthropic runs real-time safeguards on API inputs and outputs by default.”
↩︎ Key concept“There is no additional opt-in filter to enable.”
↩︎ Exam trap 1 - 2.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/mitigate-jailbreaksOfficial docs
“These attacks fall into two categories with different threat models:”
↩︎ Two threat models, two sets of controls“where the user is trusted but Claude processes third-party content”
↩︎ Two threat models, two sets of controls“Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation.”
↩︎ Input classifiers: a small model in front of the big one“Use structured outputs to constrain the response to a simple classification.”
↩︎ Input classifiers: a small model in front of the big one“Filter user input for known injection patterns before it reaches Claude.”
↩︎ Input classifiers: a small model in front of the big one“Put untrusted content only in tool results.”
↩︎ Untrusted content and the access controls behind it“Apply the principle of least privilege so that a successful injection can do minimal damage”
↩︎ Untrusted content and the access controls behind it“run tools in sandboxed environments, and scope permissions as narrowly as possible”
↩︎ Untrusted content and the access controls behind it“Red-team your own agent.”
↩︎ Untrusted content and the access controls behind it“Adjust responses and consider throttling or banning users who repeatedly attempt to circumvent your application's guardrails.”
↩︎ Tiering safeguards up to human review“instructions you place there may be ignored or flagged as a potential injection”
↩︎ Exam trap 3“Screen tool outputs before Claude acts on them.”
↩︎ Exam trap 4“Use structured outputs to constrain the response to a simple classification.”
↩︎ Exam trap 6 - 3.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusalsOfficial docs
“when streaming classifiers intervene to handle potential policy violations”
↩︎ When the platform classifier fires: stop_reason refusal“you must reset the conversation context before continuing”
↩︎ When the platform classifier fires: stop_reason refusal“you must reset the conversation context before continuing”
↩︎ Exam trap 2 - 4.
“treat any partial output as incomplete and discard it”
↩︎ When the platform classifier fires: stop_reason refusal“A refusal can arrive before any output, or mid-stream after partial output.”
↩︎ Exam trap 5 - 5.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-prompt-leakOfficial docs
“Techniques include using regular expressions, keyword filtering, or other text processing methods.”
↩︎ Output-side controls: filtering and auditability - 6.https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinationsOfficial docs
“Explicitly give Claude permission to admit uncertainty.”
↩︎ Output-side controls: filtering and auditability - 7.
“publish, reject, or send to a human reviewer”
↩︎ Tiering safeguards up to human review“Missing values evaluate to unknown”
↩︎ Tiering safeguards up to human review