What you will be able to do
- Explain why guardrails on Databricks are implemented as service policies, and how they differ from Unity Catalog grants
- Choose the right built-in guardrail (jailbreak, unsafe content, sensitive data) and phase for a malicious-input threat
- Predict what a caller receives when a guardrail blocks a request
- Identify what a built-in LLM-judge guardrail can and cannot see
Key concept
Service policy (guardrail) — A Unity Catalog policy attached to an AI service that inspects the content of each request before the model runs and each response after it, then allows, denies, or holds the interaction. On Databricks, every guardrail, built-in or custom, is a service policy.
1.Guardrails are service policies, not grants
Say your team exposes an LLM through a Model Service, and apps and agents call it. Some callers will send prompts meant to break the model's safety instructions, inject instructions of their own, or extract content you never meant to serve. The question for this objective is which control stops those inputs. On Databricks, guardrails are service policies. A service policy is attached to an AI service registered in Unity Catalog. That service can be a Databricks-hosted model or a model from an external provider such as OpenAI, Anthropic, or Google, and both are governed the same way.
Candidates often reach for Unity Catalog privileges here, so it helps to see why grants alone can't stop a malicious prompt. A grant answers one question: can this principal call this service? It is checked before the request reaches the service, and it looks only at identity and privileges. A service policy answers a different question: how must this interaction proceed? It decides per request, using the request content, the response content and the actor context. A jailbreak attempt usually comes from a principal who is allowed to call the model, so only a check on content can catch it. The two layers work together. A principal still needs the right privileges to call the service, and the service policy then checks every interaction.
A service policy is evaluated at two points. The input phase (ON CALL) runs against the request before Databricks invokes the model. The output phase (ON RESULT) runs against the response before it goes back to the caller. Defenses against malicious user input belong in the input phase, which is the place to deny a harmful prompt before it reaches a model.
Checkpoint 1 of 5· Check yourself
A user who holds EXECUTE on a Model Service keeps sending prompt-injection attempts. Which control can block those prompts?
Grants decide only whether a principal can call the service. Service policies inspect the content of each request, and the input phase runs before the model is invoked.
“Input (ON CALL): Before Databricks invokes the service, against the request.”Source: docs.databricks.com
2.Choosing a built-in guardrail
Databricks ships managed guardrails under the system.ai namespace. To use one, you select it from the Guardrail type menu with no code to write, and you need MANAGE on the target service. For malicious input, two of them matter most. system.ai.block_jailbreak denies prompt-injection and jailbreak attempts, and it runs on requests only. system.ai.block_unsafe_content denies unsafe or harmful content and can run on both phases. The other two have different jobs. system.ai.block_hallucination works on responses only, so it does nothing against a hostile prompt. system.ai.detect_sensitive_data is the Sensitive Data Detection guardrail. It is the odd one out: it is deterministic and pattern-based with no evaluator model, and it can redact a matched value instead of only blocking the request.
| Built-in policy | Denies | Phase | Decision method |
|---|---|---|---|
| system.ai.block_jailbreak | Requests that attempt to circumvent model safety instructions | Input only | LLM judge (evaluator model) |
| system.ai.block_unsafe_content | Unsafe or harmful content | Input, output or both | LLM judge (evaluator model) |
| system.ai.block_hallucination | Hallucinated responses | Output only | LLM judge (evaluator model) |
| system.ai.detect_sensitive_data | Structured sensitive data such as credit card numbers and Social Security numbers (block or redact) | Input, output or both | Deterministic pattern match, no evaluator |
The three judge-based guardrails are LLM-as-a-judge checks. Each one runs a read-only, Databricks-curated prompt against an evaluator model. A default evaluator is preselected. To use another one, open Advanced options and pick it, which needs CAN_QUERY on that model. The evaluator is separate from the service you are protecting. When a user pastes an SSN or card number into a prompt, Sensitive Data Detection on the input phase can either block the request or replace the value with a token such as [US_SSN] before the model sees it.
Checkpoint 2 of 5· Match them up
Match each built-in guardrail to its scope or behaviour
Tap a term, then the definition that fits it.
Jailbreak is input-only and hallucination is output-only. Unsafe content can run on both phases, and only Sensitive Data Detection is deterministic and can redact.
“Jailbreak (system.ai.block_jailbreak): denies prompt-injection and jailbreak attempts (requests only).”Source: docs.databricks.com
Checkpoint 3 of 5· Put it in order
Put these steps for attaching a built-in guardrail to a model service in order
- 1.Enter a name, then choose the Guardrail type (for example, Unsafe Content)
- 2.On the Models tab, select the model service
- 3.Open the Policies tab and click New policy
- 4.Set the Rank and Phase, then click Create policy
- 5.In the workspace sidebar, click AI Gateway
You reach the service through AI Gateway, open its Policies tab, create a new policy, choose the guardrail, and then set rank and phase before you create it.
“Open the Policies tab, then click New policy.”Source: docs.databricks.com
3.What the caller sees when a guardrail fires
Every service policy returns one of three outcomes. ALLOW lets the interaction proceed. DENY blocks it, and the result is not an error status. The caller gets an HTTP 200 response whose assistant turn carries a short message naming the policy that blocked, plus a top-level databricks_service_policy object with the structured detail, including the block reason. Databricks does this for conversational clients, such as coding agents, that resend the full history on every turn. If the block were an error, the client would trigger the same block again on every later turn. ASK holds the interaction for human approval. It is aimed at sensitive operations such as a destructive MCP tool call, so for a model service facing malicious prompts the outcome that matters is DENY.
Checkpoint 4 of 5· Check yourself
Why does Databricks return a DENY as a normal HTTP 200 turn instead of an error status?
Returning the block as a normal assistant turn keeps clients that resend the whole conversation, such as coding agents, from hitting the same block again on every later turn.
“keeps conversational clients that resend the full history, such as coding agents, from re-triggering the same block”Source: docs.databricks.com
A DENY comes back as a successful HTTP 200, so a status-code check never sees it. The block detail, including the reason, is in the top-level databricks_service_policy object, and the assistant turn names the policy that blocked.
Sources1
4.What a built-in judge can and cannot see
Choosing a guardrail also means knowing where it stops working. When a built-in judge runs, Databricks sends the evaluator a system message with the policy prompt and an output contract, and a user message with the content under evaluation. On the input side of a model service, that content is only the last user message. The evaluator does not see the protected service's system prompt, or any image or audio content. By default each evaluation covers one message, so an attacker who escalates slowly over several turns can get past it. On model and model provider services you can widen the input window when you attach the policy by setting how many recent turns the evaluator receives. This setting applies only to input and is not available on MCP services.
The evaluator returns flagged, an optional confidence and a reason, and a flagged verdict blocks the interaction by default. Because the judge is a model, its verdicts are non-deterministic. To audit a particular decision, enable an inference table on the evaluator model service. Built-in policies have no separate fee. Each evaluation is billed as a call to the evaluator model, so keep the number of policies per phase small and prefer a low-latency evaluator.
Checkpoint 5 of 5· Check yourself
An attacker spreads a jailbreak across six harmless-looking turns. The Jailbreak guardrail is attached with default settings and misses it. What change targets this gap?
By default the judge sees one message. On model services you can widen the input window to several recent turns. Jailbreak detection is input-only, and the evaluator never sees the system prompt.
“a built-in service policy can't detect patterns that span multiple messages, such as gradual escalation across a conversation.”Source: docs.databricks.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Selecting both phases for the Jailbreak guardrail makes it screen model responses for jailbreak content as well.Why is that wrong?
Jailbreak detection is scoped to requests. It runs only on the input phase, whatever phases you might expect it to cover.
Covered in Choosing a built-in guardrail
2.A request blocked by a guardrail comes back as an HTTP error such as 403.Why is that wrong?
A DENY returns a successful HTTP 200 response. The block detail is carried in a databricks_service_policy object.
Covered in What the caller sees when a guardrail fires
3.The jailbreak judge evaluates the prompt against the service's system prompt, so it knows what the model was told not to do.Why is that wrong?
The evaluator sees only the single extracted item. It never sees the protected service's system prompt.
Covered in What a built-in judge can and cannot see
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Service policies don't replace Unity Catalog grants.”
↩︎ Guardrails are service policies, not grants“deny a prompt that contains PII before it reaches a model.”
↩︎ Guardrails are service policies, not grants“system.ai.block_jailbreak: denies requests that attempt to circumvent model safety instructions.”
↩︎ Choosing a built-in guardrail“Unlike the others, it's deterministic (pattern-based, no evaluator model) and can redact rather than only block.”
↩︎ Choosing a built-in guardrail“ASK: the policy holds the interaction for human approval before it proceeds.”
↩︎ What the caller sees when a guardrail fires“Instead of an error status, Databricks returns a successful (HTTP 200) response.”
↩︎ What the caller sees when a guardrail fires“This applies to the input only, and isn't available on MCP services.”
↩︎ What a built-in judge can and cannot see“The evaluator is a model, so its verdicts are non-deterministic”
↩︎ What a built-in judge can and cannot see“Databricks doesn't charge a separate fee for built-in service policies.”
↩︎ What a built-in judge can and cannot see“Service policies are the mechanisms you use to implement guardrails for AI services.”
↩︎ Key concept“jailbreak detection runs only on the input phase and hallucination detection only on the output phase”
↩︎ Exam trap 1“Instead of an error status, Databricks returns a successful (HTTP 200) response.”
↩︎ Exam trap 2“It doesn't see the system prompt of the protected service or image and audio content.”
↩︎ Exam trap 3“Input (ON CALL): Before Databricks invokes the service, against the request.”
↩︎ Checkpoint“jailbreak detection runs only on the input phase and hallucination detection only on the output phase”
↩︎ Prediction“keeps conversational clients that resend the full history, such as coding agents, from re-triggering the same block”
↩︎ Checkpoint“a built-in service policy can't detect patterns that span multiple messages, such as gradual escalation across a conversation.”
↩︎ Checkpoint - 2.
“On Databricks, guardrails are service policies.”
↩︎ Guardrails are service policies, not grants“Jailbreak (system.ai.block_jailbreak): denies prompt-injection and jailbreak attempts (requests only).”
↩︎ Checkpoint“Open the Policies tab, then click New policy.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/data-governance/unity-catalog/service-policies/detect-sensitive-dataOfficial docs
“Databricks replaces each matched value with a placeholder token, such as [US_SSN], and forwards the rewritten content.”
↩︎ Choosing a built-in guardrail