What you will be able to do
- Explain how a Unity Catalog column mask works as a guardrail that keeps sensitive values out of GenAI data pipelines
- Choose between ABAC policies, table-level column masks and dynamic views based on governance and performance needs
- Design masking UDFs and policies that meet a query performance objective
- Spot the places in a GenAI data flow where source-table masks are not enforced, such as AI Search indexes and pipeline refreshes
Key concept
Column mask — A column mask is a SQL function that Unity Catalog applies to a column every time it is queried. Depending on who is asking, it returns either the real value or a redacted one. The guardrail sits in the data layer itself, so every consumer gets the same protection, including a GenAI app that reads the table.
1.Masking as a guardrail in the data layer
A GenAI application only knows what its data tells it. If a retrieval pipeline, a fine-tuning set or an agent tool can read raw social security numbers, emails or salaries, that data can turn up in a prompt, a retrieved chunk or a generated answer. A guardrail is any control that stops this from happening. Masking is the guardrail that works at the data itself rather than on the model's input or output.
In Unity Catalog, masking is done with column masks. A column mask is a SQL user-defined function bound to a column. At query time it gets the column value and returns either that value or a substitute, such as ***, the last four digits, or a hashed placeholder. One rule constrains every mask: the value it returns has to fit the column. A string such as 'CONFIDENTIAL' can't be returned into a DOUBLE column. Each column can have only one mask.
Masks can be attached in two ways. A table owner can bind one directly with ALTER TABLE ... ALTER COLUMN ... SET MASK, which applies to just that column on that table. Alternatively, an attribute-based access control (ABAC) policy can find columns by their governed tags, such as pii, and mask every matching column across a catalog or schema. The policy below masks every pii-tagged column in a catalog for all users. The only exception is anyone who has assumed a designated PII-cleared role.
CREATE POLICY pii_default_mask
ON CATALOG customer_data
COLUMN MASK mask_pii
TO `account users`
EXCEPT `role-pii-cleared`
FOR TABLES
MATCH COLUMNS has_tag('pii') AS pii_col
ON COLUMN pii_col;Checkpoint 1 of 6· Fill the gap
This policy masks SSNs for US analysts, but admins should see raw values. Which keyword completes it?
CREATE POLICY mask_ssn
ON SCHEMA prod.customers
COLUMN MASK ssn_to_last_nr
TO us_analysts ? admins
FOR TABLES
MATCH COLUMNS has_tag_value('pii', 'ssn') AS ssn
ON COLUMN ssn
USING COLUMNS (4);EXCEPT lists the principals that are exempt from the policy. WHEN filters tables by their tags, and USING COLUMNS passes arguments to the UDF.
Source: docs.databricks.comSources1
2.Choosing the layer: ABAC, table-level masks or dynamic views
A performance objective is only one constraint. You also have to decide who owns the rule and whether a determined user can get around it. Unity Catalog gives you three options, and each makes a different trade.
Dynamic views put masking logic, such as CASE WHEN is_account_group_member('auditors') THEN email ELSE ... END, inside a view definition. They are the fastest of the three, but they have two governance weaknesses. They carry no tags or policy metadata in system tables, which makes them hard to audit at scale. They also lack a SecureView barrier, so a user can write a predicate with side effects and use it to infer the values of filtered rows.
Table-level column masks are set per table by its owner. That suits a small, stable set of tables with unusual logic, but the owner can also remove the mask. ABAC policies attach at the catalog or schema level, match columns by governed tag and cover new tagged tables automatically. Table owners cannot override them. Databricks recommends ABAC for consistent PII masking, and the docs also note that ABAC policy logic is evaluated more efficiently than table-specific UDFs.
| Approach | Defined with | Who controls it | Performance / security note |
|---|---|---|---|
| Dynamic views | SQL logic in the view definition | View author | Full query optimization and predicate pushdown, but no SecureView barrier against probing |
| Table-level column masks | ALTER TABLE ... ALTER COLUMN ... SET MASK | Table owner, who can modify or remove it | Each table must be configured individually |
| ABAC policies | CREATE POLICY ... ON CATALOG/SCHEMA/TABLE | Catalog or schema owners; table owners cannot override | Matched dynamically by governed tags using has_tag() and has_tag_value() |
Checkpoint 2 of 6· Check yourself
A team picks a dynamic view to mask PII because it gives the best query speed. Which governance risk are they accepting?
Dynamic views have no SecureView barrier, which leaves them open to probing attacks. Row filters and column masks are protected against this.
“Because they lack a SecureView barrier, they don't protect against probing attacks”Source: docs.databricks.com
3.Designing masks that meet a performance objective
Masking always costs something. Column masks and row filters make sure users never see base values before the protection is applied. When security and speed conflict, the engine chooses security. It will give up an optimization rather than risk exposing a masked value. To meet a latency or throughput target, you have to design the mask so the engine rarely faces that choice.
The guidance comes down to a short checklist:
- Keep UDFs simple. Use basic CASE expressions and built-in SQL functions. Avoid subqueries, joins against large tables, external API calls and per-row metadata lookups.
- Prefer SQL UDFs over Python UDFs. Python offers fewer optimization opportunities.
- Use deterministic expressions that can't throw errors. An expression like ANSI division can fail with a 'division by zero' error that reveals something about the value, so the compiler won't push it down. Use try_divide instead.
- Mask only truly sensitive columns and reuse functions. Each distinct mask on a table adds overhead.
- Avoid regex masking on large text fields. Running a regex over serialized documents makes the engine rewrite the whole payload on every row. GenAI tables full of raw document text are especially exposed to this.
- Target principals with TO/EXCEPT. When a principal is exempt, the UDF doesn't run for them at all. If you do use identity functions such as is_account_group_member() inside a UDF, they are resolved once during query analysis, not per row.
Checkpoint 3 of 6· Check yourself
A mask UDF calls is_account_group_member() three times with different groups. What does this cost per row on a billion-row table?
Identity functions are resolved once during query analysis, and several group checks are combined into a single UC API call. The per-row cost is therefore minimal.
“These functions are resolved once during query analysis, not per row.”Source: docs.databricks.com
Checkpoint 4 of 6· Match them up
Match each design choice to its performance effect
Tap a term, then the definition that fits it.
Each of these follows from the performance guidance. Exemptions skip the UDF entirely, regex on large text is expensive, distinct masks add overhead, and expressions that can throw errors limit pushdown.
“The EXCEPT clause eliminates the policy entirely for exempt users, which means no UDF execution for those users.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A customer-support chatbot serving latency-sensitive users occasionally receives inputs containing phone numbers or email addresses that must never reach the underlying LLM. The team also does not want legitimate support requests to fail outright just because they mention contact details. Which AI Gateway guardrail configuration meets this objective?
Correct answer: A — Configure a PII sanitizing guardrail that replaces detected phone numbers and email addresses with placeholder tokens before the request reaches the model.
- A. A sanitizing PII guardrail redacts the sensitive values with placeholder tokens before the model ever sees them, letting the request continue to succeed while ensuring the raw phone number or email never reaches the LLM.
- B. A blocking PII guardrail terminates the request with an error instead of letting it proceed, which fails the requirement that legitimate support requests should not be rejected outright.
- C. Hallucination-blocking guardrails target fabricated facts or invented citations in model output, not real PII supplied by the user in the request, so it does not address the stated requirement.
- D. A custom prompt-based instruction relies on the model choosing to comply and still exposes the raw phone number or email to the model itself, which does not prevent the sensitive data from reaching the LLM.
4.Where source-table masks stop: indexes, pipelines and requests
No. A mask protects queries against the table. It does not protect copies of the data made elsewhere. An AI Search index built from an ABAC-protected table syncs every row and enforces no row filter or column mask when it serves queries. The guardrail is to keep masked columns out of the index altogether, using the columns-to-sync setting.
Pipelines have a related gap. When a pipeline refreshes a materialized view or streaming table, policies are evaluated against the pipeline owner or run-as identity. If that identity is subject to a mask, the output permanently holds masked data. The fix is to exempt the pipeline identity with EXCEPT. Keep in mind that exempt principals see completely unmasked data, so only trusted service principals should be exempted.
Masking also happens at the request boundary of a served model. Databricks' Sensitive Data Detection is tuned separately for its two actions. Redaction is tuned for recall, because missing a value is worse than over-redacting. Blocking is tuned for precision. According to the documentation, detection adds well under 50 ms to a request, so it can sit in a latency-sensitive path.
Checkpoint 6 of 6· Check yourself
A pipeline that builds a feature table runs as a service principal that falls under a PII mask policy. What happens to the materialized view?
Pipelines evaluate policies as the run-as identity. If that identity is masked, the masked values are written into the output.
“the materialized view or streaming table permanently contains masked or filtered data”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Dynamic views are the safest way to mask PII because they are the fastest and the most flexible.Why is that wrong?
Dynamic views are fast but have no SecureView barrier, so users can probe them with crafted predicates. They are also harder to audit at scale.
Covered in Choosing the layer: ABAC, table-level masks or dynamic views
2.Masking the source table automatically protects the AI Search index built from it.Why is that wrong?
Indexes sync all rows and do not enforce masks when serving queries. Leave masked columns out of the index using the columns-to-sync setting.
Covered in Where source-table masks stop: indexes, pipelines and requests
3.To be safe, put a distinct mask on every column of a large table.Why is that wrong?
Each distinct mask adds query overhead. Mask only truly sensitive columns, and reuse one function across columns where you can.
Covered in Designing masks that meet a performance objective
4.A mask can return any placeholder string, such as 'CONFIDENTIAL', for any column.Why is that wrong?
The mask's return value must be compatible with the column's type, so a numeric column needs a numeric placeholder.
Covered in Masking as a guardrail in the data layer
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The mask is a SQL UDF that takes the column value as input and returns the original value or a masked version.”
↩︎ Masking as a guardrail in the data layer“Policy logic is also evaluated more efficiently than table-specific UDFs.”
↩︎ Choosing the layer: ABAC, table-level masks or dynamic views“When the query engine must choose between optimization and protecting against information leakage from filtered or masked values, it always makes the secure choice”
↩︎ Designing masks that meet a performance objective“Expressions that can throw errors (such as ANSI division) prevent the SQL compiler from pushing operations down in the query plan”
↩︎ Designing masks that meet a performance objective“The mask is a SQL UDF that takes the column value as input and returns the original value or a masked version.”
↩︎ Key concept“The return type must match or be castable to the column's data type.”
↩︎ Exam trap 4 - 2.
“individual table owners can't remove, modify, or bypass it”
↩︎ Choosing the layer: ABAC, table-level masks or dynamic views“Because they lack a SecureView barrier, they don't protect against probing attacks”
↩︎ Exam trap 1“Dynamic views fully support query optimization and predicate pushdown, so they can offer better query performance than row filters and column masks.”
↩︎ Prediction“Because they lack a SecureView barrier, they don't protect against probing attacks”
↩︎ Checkpoint - 3.
“Regex-based masking on serialized documents forces the engine to scan and rewrite the entire payload for every row.”
↩︎ Designing masks that meet a performance objective“Each distinct mask on a table adds overhead; reusing the same function across columns can reduce it.”
↩︎ Exam trap 3“These functions are resolved once during query analysis, not per row.”
↩︎ Checkpoint“The EXCEPT clause eliminates the policy entirely for exempt users, which means no UDF execution for those users.”
↩︎ Checkpoint - 4.
“For tables with column masks, you can exclude masked columns from the index using the columns to sync setting.”
↩︎ Where source-table masks stop: indexes, pipelines and requests“Exempted principals see unfiltered, unmasked data.”
↩︎ Where source-table masks stop: indexes, pipelines and requests“ABAC policies on a source table don't apply to AI Search indexes created from that table.”
↩︎ Exam trap 2“the materialized view or streaming table permanently contains masked or filtered data”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/data-governance/unity-catalog/service-policies/detect-sensitive-dataOfficial docs
“Redaction should catch as much sensitive data as possible, since a missed value is worse than an over-redaction, so it is optimized for recall.”
↩︎ Where source-table masks stop: indexes, pipelines and requests“Detection is deterministic and fast. It adds well under 50 ms to a request, even on a large (100-turn) conversation.”
↩︎ Where source-table masks stop: indexes, pipelines and requests