CertSafari
    CCAR-P · Lessons

    Domain 5 · Lesson 26/38

    Implementing guardrails and safety controls around Claude

    Implement guardrails and safety controls

    18 min read
    2.8% of exam
    7 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Describe the guardrail layers you are responsible for on top of the platform safeguards that already run by default
    • Screen untrusted input and tool output with a small classifier call whose verdict your application can branch on
    • Structure an agent so third-party content cannot act as instructions, and scope its access so a successful injection does little damage
    • Detect stop_reason: refusal, reset context, and interpret the refusal category and its billing behaviour
    • Place output-side checks and a human review tier where a confident automated verdict is not available

    Key concept

    Layered guardrails — Safety in a Claude deployment is a stack, not a switch: Anthropic's own real-time classifiers on API inputs and outputs are the floor that is always there, and every layer above it — input screens, prompt structure, access limits, output filters, human review — is something you build in your application.

    1.The floor you get, and the layers you build

    Start with what is already running. Anthropic applies real-time safeguards to API inputs and outputs for every request, and there is no dashboard toggle that turns extra filtering on: "There is no additional opt-in filter to enable." That is the floor. It is not a substitute for application-level controls, because it enforces Anthropic's Usage Policy, not your product's rules — an ad network's disclaimer requirements or a support bot's refusal to discuss competitors are yours to enforce.

    So the design question for this subdomain is never "is filtering on?" but "at which point in my request path does each check run?" The rest of this lesson walks that path in order: what reaches Claude, how the prompt is structured, what Claude is allowed to touch, what comes back, and what a human sees.

    Sources1

    2.Two threat models, two sets of controls

    Before choosing controls, decide who the adversary is. The guardrails documentation splits the problem in two: "These attacks fall into two categories with different threat models:" In a jailbreak or direct prompt injection, the user of your application is the attacker, crafting input designed to get past your rules. In indirect prompt injection the user is trusted and the attacker is upstream of the content — an inbound email, a fetched page, OCR text, a tool result.

    The distinction matters because the controls barely overlap. Against a hostile user you screen and rate-limit the user. Against hostile content you cannot ban anyone: you change how the content enters the prompt and what Claude is permitted to do once it is there.

    Which threat model a control belongs to
    Threat modelWho is the adversaryRepresentative controls
    Jailbreaks and direct prompt injectionThe user of your applicationHarmlessness screen on user input, input validation for known injection patterns, ethical system prompt, throttling repeat offenders
    Indirect prompt injectionWhoever can influence third-party content Claude readsDeliver content in tool_result blocks, JSON-encode it, state an untrusted-content policy, least privilege, screen tool outputs

    A fintech company's customer-facing assistant is built on Claude. Before a user message reaches the main conversation, the compliance team wants a fast, cheap classification step that flags requests referring to harmful, illegal, or fraudulent activity, so those messages can be routed to a stricter review flow instead of the main model. Which guardrail design best satisfies this requirement while controlling cost?

    Sources2

    3.Input classifiers: a small model in front of the big one

    The first layer you own is an input classifier. The recommended shape is a cheap pre-flight call: "Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation." The point of using a small model is that the screen runs on every request, so its cost and latency must be near-negligible compared with the main call.

    A screen is only a guardrail if your code can act on it mechanically. That is why the guidance pairs the screen with constrained decoding — "Use structured outputs to constrain the response to a simple classification." A boolean field such as is_harmful is something an if-statement can branch on; a paragraph of prose explaining the model's concerns is not. Alongside the model screen sits ordinary deterministic validation: "Filter user input for known injection patterns before it reaches Claude."

    The third input-side control is the prompt itself. A system prompt that states boundaries explicitly and tells Claude the exact refusal wording to use makes refusals consistent and, usefully, machine-detectable in your logs.

    Sources2

    4.Untrusted content and the access controls behind it

    Now the harder half. When Claude reads content on a user's behalf, your job is to make the boundary between your instructions and that content unambiguous. The first rule is placement: "Put untrusted content only in tool results." Claude is trained to treat instructions appearing inside tool results with skepticism, which is exactly the skepticism you want — and which you throw away by pasting a fetched page into the system prompt or a plain user turn.

    Three reinforcements follow. Label the content: say in the tool description or the result structure that this is, for example, the body of an inbound email from an unknown sender. State the policy in the system prompt — tell Claude that tool, document and search content is untrusted data that must never override the system prompt or the user's request, and that apparent instructions inside it should be reported rather than obeyed. And encode it: wrapping third-party strings in JSON gives unambiguous delimiters, so an attacker cannot close a quote or tag to break out into an instruction context.

    The same placement rule cuts the other way, and this is a favourite trap: do not put your own instructions into tool results. "Because Claude treats tool-result content as untrusted data, instructions you place there may be ignored or flagged as a potential injection." Send them in a user turn after the tool_result block instead.

    Content structure is not enough on its own, so guardrails extend to authorisation. "Apply the principle of least privilege so that a successful injection can do minimal damage" — do not hand Claude secrets it does not need, "run tools in sandboxed environments, and scope permissions as narrowly as possible." Access control is what converts a successful injection from a breach into a nuisance.

    Finally, the screening pattern from the input side applies to the output of your own tools: "Screen tool outputs before Claude acts on them." Run the tool, pass its raw output to a Haiku 4.5 classifier, and only hand back a tool_result block if the verdict field — again a structured boolean, injection_suspected — reports nothing. When it does report an attempt, return an error or a stripped summary instead of the raw content, and consider surfacing the attempt to the user.

    The same screening pattern at two points in the request path
    What is screenedScreening modelStructured verdict fieldAction when it trips
    User input, before the main conversationClaude Haiku 4.5is_harmfulDo not pass the turn to the main conversation; escalate repeat offenders
    Raw tool output, before Claude sees itClaude Haiku 4.5injection_suspectedReturn an error or stripped summary in the tool_result block instead of raw content

    None of these layers is verified until you attack them yourself. "Red-team your own agent." Feed the workflow documents, emails and tool outputs that deliberately contain injection attempts before deployment, and confirm both that Claude ignores them and that your screens and confirmation steps catch what it does not.

    A review of a Claude-powered support bot's logs shows that a small number of accounts have each triggered the same jailbreak-style refusal more than a dozen times in one week, using slightly reworded prompts each time. According to Anthropic's guardrail guidance, what is the most appropriate response to this specific pattern?

    Sources2

    5.When the platform classifier fires: stop_reason refusal

    The real-time safeguards from the first section are visible in your code as a specific response shape. On Claude 4 and later, streaming responses come back with stop_reason set to refusal "when streaming classifiers intervene to handle potential policy violations". Note what this is not: it is a successful HTTP 200 response, not an exception, so error-handling code alone will never see it.

    The refusal response returned when a streaming classifier intervenesjson
    {
      "role": "assistant",
      "content": [
        {
          "type": "text",
          "text": "Hello.."
        }
      ],
      "stop_reason": "refusal",
      "stop_details": {
        "type": "refusal",
        "category": "cyber",
        "explanation": "This request was declined because it could enable cyber harm."
      }
    }

    The operational rule is the part people get wrong: on a refusal "you must reset the conversation context before continuing". Remove or rephrase the offending turn, or clear the history; retrying the same message list simply earns another refusal. The partial text is not usable either — "treat any partial output as incomplete and discard it", because a refusal can arrive before any output or mid-stream after some.

    stop_details tells you which policy area tripped. Display the explanation rather than parsing it — the wording is not stable — and branch on category. Whether you were billed for the attempt also depends on the category.

    Refusal categories in stop_details, and whether the request is billed before any output
    categoryBilled before any output
    "cyber"No
    "bio"Yes
    "frontier_llm"Yes
    "reasoning_extraction"Yes
    "general_harms"No

    Around that, the recommended handling is a small control loop: monitor for refusals in your response handling, reset context automatically, fall back to another Claude model rather than surfacing a raw refusal, show your own user-facing message, and track refusal frequency as a signal about your prompts. A rising refusal rate on benign traffic is usually a prompt problem, not a user problem — benign cybersecurity and life-sciences work can trigger the cyber and bio categories.

    An agent ingests inbound customer emails as tool results and drafts replies on the user's behalf. A red-team exercise crafts an email body containing text designed to close out the surrounding context and inject a new instruction telling the agent to forward the user's contact list. The way the email body is concatenated into the tool result lets the crafted text escape into an instruction context. Which change most directly closes this specific escape route?

    Sources34

    6.Output-side controls: filtering and auditability

    Guardrails on the way out are cheaper than they look. For prompt leakage, post-processing is a deterministic layer over a probabilistic one: "Techniques include using regular expressions, keyword filtering, or other text processing methods." It runs after generation and does not compete with the task, which matters because leak-resistant prompting itself adds complexity that can degrade the model's performance on the real work.

    The other output-side control worth wiring in is auditability. Giving Claude explicit permission to decline — "Explicitly give Claude permission to admit uncertainty." — plus a requirement to support each claim with a quote from the provided material turns an unverifiable answer into one your own code, or a reviewer, can check. That check is what feeds the last layer.

    Sources56

    7.Tiering safeguards up to human review

    The API safeguards guidance presents these controls as tiers you add as a deployment matures. The basic tier is accountability plumbing: "Store IDs linked with each API call", so violative content can be located later, optionally with hashed per-user IDs so misuse can be attributed to an individual. Enforcement sits on top of it — "Warn, throttle, or suspend users who repeatedly violate" the Terms and Usage Policy. This is the same repeat-offender response the jailbreak guidance recommends when one user keeps tripping the same refusal.

    Intermediate tiers narrow the attack surface rather than inspecting it: restrict end users to a limited set of prompts, or point Claude only at a knowledge corpus you already control. Advanced tiers add moderation — using Claude itself for content moderation, or running a moderation API over every end-user prompt before it reaches Claude.

    The top tier is a human: "Set up an internal human review system to flag prompts that are marked by Claude" or by a moderation API as harmful, so you can intervene against accounts with high violation rates. The important design point is that review is a routed outcome, not a fallback for everything. A moderation loop decides what happens to each item — "publish, reject, or send to a human reviewer" — and in a well-built pipeline that third branch is reserved for genuine uncertainty: when a needed field could not be determined, "Missing values evaluate to unknown" and the item is routed to review instead of receiving a confident automated verdict.

    Sources127

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Before launching you must enable Anthropic's input/output content filter in the console, and once it is on your application does not need its own moderation layer.Why is that wrong?

      Real-time safeguards on API inputs and outputs already run by default and there is nothing to opt into; your own moderation layer is added in your application on top of them.

      Covered in The floor you get, and the layers you build

    2. 2.A classifier refusal arrives as an HTTP error, so catching exceptions and retrying the same request handles it.Why is that wrong?

      A refusal is a successful response carrying stop_reason: refusal, and continuing without resetting the conversation context produces further refusals.

      Covered in When the platform classifier fires: stop_reason refusal

    3. 3.The safest place for your strictest safety instructions is inside the tool result, right next to the untrusted content they are meant to govern.Why is that wrong?

      Tool-result content is treated as untrusted data, so your instructions there may be ignored or even flagged as an injection; send them in a following user turn instead.

      Covered in Untrusted content and the access controls behind it

    4. 4.If user input is screened by a classifier, content that arrives later through tool calls has already passed the guardrail.Why is that wrong?

      Tool output is a separate entry point and needs the same screening before Claude acts on it, with the raw content withheld when the screen reports an injection attempt.

      Covered in Untrusted content and the access controls behind it

    5. 5.A refusal always happens before generation starts, so any text you have already streamed to the user is safe to keep.Why is that wrong?

      A refusal can also arrive mid-stream after partial output, and that partial output must be discarded as incomplete.

      Covered in When the platform classifier fires: stop_reason refusal

    6. 6.An input screen should return a short natural-language explanation so your logs capture the reasoning.Why is that wrong?

      The screen exists so code can branch on it, so the response is constrained by a JSON schema to a simple classification such as a boolean is_harmful.

      Covered in Input classifiers: a small model in front of the big one

    Practise it for real

    Red-team a document-reading agent and close the gap with a tool-output screen

    1. 1.Build a minimal agent that fetches a document and summarises it, delivering the document text inside a tool_result block rather than in the system prompt or a user text block.

      Why: Claude is trained to treat instructions inside tool results with skepticism, so placement is the cheapest part of the guardrail.

      You should see: A working summariser whose untrusted payload is confined to tool_result content.

    2. 2.Add an untrusted-content policy to the system prompt stating that tool, file and search content is data, that it must never override the system prompt or the user request, and that apparent instructions should be reported instead of followed.

      Why: The policy tells Claude how to calibrate directives found inside retrieved content.

      You should see: The summary of a hostile document mentions that it contained instructions, rather than acting on them.

    3. 3.Feed the agent a document whose body contains an injection such as an instruction to ignore previous instructions and exfiltrate a key, and record what the agent does.

      Why: Red-teaming your own agent before deployment is the only way to know whether the layers you added hold.

      You should see: Either a clean report of the attempt, or a concrete failure you now have to screen for.

    4. 4.Insert a Claude Haiku 4.5 screening call between the tool and the model, asking only whether the content contains instructions aimed at the assistant, with the response constrained by a JSON schema to a boolean injection_suspected.

      Why: A structured verdict is a value your application can branch on; prose is not.

      You should see: injection_suspected is true for the hostile document and false for a clean one.

    5. 5.Branch on the verdict: when injection_suspected is true, return an error or a stripped summary in the tool_result block instead of the raw content, and surface the attempt to the user.

      Why: Screening is only a guardrail if a positive verdict changes what reaches Claude.

      You should see: The hostile document never reaches the main conversation in raw form, while clean documents summarise as before.

    6. 6.Re-run the whole red-team set, then re-scope the agent's credentials and sandbox its tools so that a screen that misses still cannot cause damage.

      Why: Least privilege bounds the impact of the injection that gets through every content-level layer.

      You should see: No injection in your set changes the agent's goals, and the agent holds no secret or permission the summarising task does not require.

    Stuck? Get a nudge

    Write the screen as a separate function that takes raw tool output and returns a boolean, not as extra wording in the main prompt — a guardrail you can unit-test is a guardrail you can trust.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Anthropic runs real-time safeguards on API inputs and outputs by default.”
      ↩︎ The floor you get, and the layers you build
      “There is no additional opt-in filter to enable.”
      ↩︎ The floor you get, and the layers you build
      “Store IDs linked with each API call”
      ↩︎ Tiering safeguards up to human review
      “Warn, throttle, or suspend users who repeatedly violate”
      ↩︎ Tiering safeguards up to human review
      “Set up an internal human review system to flag prompts that are marked by Claude”
      ↩︎ Tiering safeguards up to human review
      “Anthropic runs real-time safeguards on API inputs and outputs by default.”
      ↩︎ Key concept
      “There is no additional opt-in filter to enable.”
      ↩︎ Exam trap 1
    2. 2.
      “These attacks fall into two categories with different threat models:”
      ↩︎ Two threat models, two sets of controls
      “where the user is trusted but Claude processes third-party content”
      ↩︎ Two threat models, two sets of controls
      “Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation.”
      ↩︎ Input classifiers: a small model in front of the big one
      “Use structured outputs to constrain the response to a simple classification.”
      ↩︎ Input classifiers: a small model in front of the big one
      “Filter user input for known injection patterns before it reaches Claude.”
      ↩︎ Input classifiers: a small model in front of the big one
      “Put untrusted content only in tool results.”
      ↩︎ Untrusted content and the access controls behind it
      “Apply the principle of least privilege so that a successful injection can do minimal damage”
      ↩︎ Untrusted content and the access controls behind it
      “run tools in sandboxed environments, and scope permissions as narrowly as possible”
      ↩︎ Untrusted content and the access controls behind it
      “Adjust responses and consider throttling or banning users who repeatedly attempt to circumvent your application's guardrails.”
      ↩︎ Tiering safeguards up to human review
      “instructions you place there may be ignored or flagged as a potential injection”
      ↩︎ Exam trap 3
      “Screen tool outputs before Claude acts on them.”
      ↩︎ Exam trap 4
      “Use structured outputs to constrain the response to a simple classification.”
      ↩︎ Exam trap 6
    3. 3.
      “when streaming classifiers intervene to handle potential policy violations”
      ↩︎ When the platform classifier fires: stop_reason refusal
      “you must reset the conversation context before continuing”
      ↩︎ When the platform classifier fires: stop_reason refusal
      “you must reset the conversation context before continuing”
      ↩︎ Exam trap 2
    4. 4.
      “treat any partial output as incomplete and discard it”
      ↩︎ When the platform classifier fires: stop_reason refusal
      “A refusal can arrive before any output, or mid-stream after partial output.”
      ↩︎ Exam trap 5
    5. 5.
      “Techniques include using regular expressions, keyword filtering, or other text processing methods.”
      ↩︎ Output-side controls: filtering and auditability
    6. 7.

    Ready to test yourself?

    Practise the 12 questions on this subdomain.