CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 5 · Lesson 43/56

    Column Masks as Guardrails: Masking PII Without Losing Performance

    Use masking techniques as guard rails to meet a performance objective

    13 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain how a Unity Catalog column mask works as a guardrail that keeps sensitive values out of GenAI data pipelines
    • Choose between ABAC policies, table-level column masks and dynamic views based on governance and performance needs
    • Design masking UDFs and policies that meet a query performance objective
    • Spot the places in a GenAI data flow where source-table masks are not enforced, such as AI Search indexes and pipeline refreshes

    Key concept

    Column mask — A column mask is a SQL function that Unity Catalog applies to a column every time it is queried. Depending on who is asking, it returns either the real value or a redacted one. The guardrail sits in the data layer itself, so every consumer gets the same protection, including a GenAI app that reads the table.

    1.Masking as a guardrail in the data layer

    A GenAI application only knows what its data tells it. If a retrieval pipeline, a fine-tuning set or an agent tool can read raw social security numbers, emails or salaries, that data can turn up in a prompt, a retrieved chunk or a generated answer. A guardrail is any control that stops this from happening. Masking is the guardrail that works at the data itself rather than on the model's input or output.

    In Unity Catalog, masking is done with column masks. A column mask is a SQL user-defined function bound to a column. At query time it gets the column value and returns either that value or a substitute, such as ***, the last four digits, or a hashed placeholder. One rule constrains every mask: the value it returns has to fit the column. A string such as 'CONFIDENTIAL' can't be returned into a DOUBLE column. Each column can have only one mask.

    Masks can be attached in two ways. A table owner can bind one directly with ALTER TABLE ... ALTER COLUMN ... SET MASK, which applies to just that column on that table. Alternatively, an attribute-based access control (ABAC) policy can find columns by their governed tags, such as pii, and mask every matching column across a catalog or schema. The policy below masks every pii-tagged column in a catalog for all users. The only exception is anyone who has assumed a designated PII-cleared role.

    ABAC column mask policy: mask every column tagged pii across the customer_data catalog, except for the PII-cleared rolesql
    CREATE POLICY pii_default_mask
    ON CATALOG customer_data
    COLUMN MASK mask_pii
    TO `account users`
    EXCEPT `role-pii-cleared`
    FOR TABLES
    MATCH COLUMNS has_tag('pii') AS pii_col
    ON COLUMN pii_col;

    Checkpoint 1 of 6· Fill the gap

    This policy masks SSNs for US analysts, but admins should see raw values. Which keyword completes it?

    CREATE POLICY mask_ssn
    ON SCHEMA prod.customers
    COLUMN MASK ssn_to_last_nr
    TO us_analysts  ?  admins
    FOR TABLES
    MATCH COLUMNS has_tag_value('pii', 'ssn') AS ssn
    ON COLUMN ssn
    USING COLUMNS (4);

    Sources1

    2.Choosing the layer: ABAC, table-level masks or dynamic views

    A performance objective is only one constraint. You also have to decide who owns the rule and whether a determined user can get around it. Unity Catalog gives you three options, and each makes a different trade.

    Dynamic views put masking logic, such as CASE WHEN is_account_group_member('auditors') THEN email ELSE ... END, inside a view definition. They are the fastest of the three, but they have two governance weaknesses. They carry no tags or policy metadata in system tables, which makes them hard to audit at scale. They also lack a SecureView barrier, so a user can write a predicate with side effects and use it to infer the values of filtered rows.

    Table-level column masks are set per table by its owner. That suits a small, stable set of tables with unusual logic, but the owner can also remove the mask. ABAC policies attach at the catalog or schema level, match columns by governed tag and cover new tagged tables automatically. Table owners cannot override them. Databricks recommends ABAC for consistent PII masking, and the docs also note that ABAC policy logic is evaluated more efficiently than table-specific UDFs.

    Where the masking logic lives, and what each choice trades
    ApproachDefined withWho controls itPerformance / security note
    Dynamic viewsSQL logic in the view definitionView authorFull query optimization and predicate pushdown, but no SecureView barrier against probing
    Table-level column masksALTER TABLE ... ALTER COLUMN ... SET MASKTable owner, who can modify or remove itEach table must be configured individually
    ABAC policiesCREATE POLICY ... ON CATALOG/SCHEMA/TABLECatalog or schema owners; table owners cannot overrideMatched dynamically by governed tags using has_tag() and has_tag_value()

    Checkpoint 2 of 6· Check yourself

    A team picks a dynamic view to mask PII because it gives the best query speed. Which governance risk are they accepting?

    Sources21

    3.Designing masks that meet a performance objective

    Masking always costs something. Column masks and row filters make sure users never see base values before the protection is applied. When security and speed conflict, the engine chooses security. It will give up an optimization rather than risk exposing a masked value. To meet a latency or throughput target, you have to design the mask so the engine rarely faces that choice.

    The guidance comes down to a short checklist:

    - Keep UDFs simple. Use basic CASE expressions and built-in SQL functions. Avoid subqueries, joins against large tables, external API calls and per-row metadata lookups. - Prefer SQL UDFs over Python UDFs. Python offers fewer optimization opportunities. - Use deterministic expressions that can't throw errors. An expression like ANSI division can fail with a 'division by zero' error that reveals something about the value, so the compiler won't push it down. Use try_divide instead. - Mask only truly sensitive columns and reuse functions. Each distinct mask on a table adds overhead. - Avoid regex masking on large text fields. Running a regex over serialized documents makes the engine rewrite the whole payload on every row. GenAI tables full of raw document text are especially exposed to this. - Target principals with TO/EXCEPT. When a principal is exempt, the UDF doesn't run for them at all. If you do use identity functions such as is_account_group_member() inside a UDF, they are resolved once during query analysis, not per row.

    Checkpoint 3 of 6· Check yourself

    A mask UDF calls is_account_group_member() three times with different groups. What does this cost per row on a billion-row table?

    Checkpoint 4 of 6· Match them up

    Match each design choice to its performance effect

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 6· Exam question

    A customer-support chatbot serving latency-sensitive users occasionally receives inputs containing phone numbers or email addresses that must never reach the underlying LLM. The team also does not want legitimate support requests to fail outright just because they mention contact details. Which AI Gateway guardrail configuration meets this objective?

    Sources13

    4.Where source-table masks stop: indexes, pipelines and requests

    No. A mask protects queries against the table. It does not protect copies of the data made elsewhere. An AI Search index built from an ABAC-protected table syncs every row and enforces no row filter or column mask when it serves queries. The guardrail is to keep masked columns out of the index altogether, using the columns-to-sync setting.

    Pipelines have a related gap. When a pipeline refreshes a materialized view or streaming table, policies are evaluated against the pipeline owner or run-as identity. If that identity is subject to a mask, the output permanently holds masked data. The fix is to exempt the pipeline identity with EXCEPT. Keep in mind that exempt principals see completely unmasked data, so only trusted service principals should be exempted.

    Masking also happens at the request boundary of a served model. Databricks' Sensitive Data Detection is tuned separately for its two actions. Redaction is tuned for recall, because missing a value is worse than over-redacting. Blocking is tuned for precision. According to the documentation, detection adds well under 50 ms to a request, so it can sit in a latency-sensitive path.

    Checkpoint 6 of 6· Check yourself

    A pipeline that builds a feature table runs as a service principal that falls under a PII mask policy. What happens to the materialized view?

    Sources45

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Dynamic views are the safest way to mask PII because they are the fastest and the most flexible.Why is that wrong?

      Dynamic views are fast but have no SecureView barrier, so users can probe them with crafted predicates. They are also harder to audit at scale.

      Covered in Choosing the layer: ABAC, table-level masks or dynamic views

    2. 2.Masking the source table automatically protects the AI Search index built from it.Why is that wrong?

      Indexes sync all rows and do not enforce masks when serving queries. Leave masked columns out of the index using the columns-to-sync setting.

      Covered in Where source-table masks stop: indexes, pipelines and requests

    3. 3.To be safe, put a distinct mask on every column of a large table.Why is that wrong?

      Each distinct mask adds query overhead. Mask only truly sensitive columns, and reuse one function across columns where you can.

      Covered in Designing masks that meet a performance objective

    4. 4.A mask can return any placeholder string, such as 'CONFIDENTIAL', for any column.Why is that wrong?

      The mask's return value must be compatible with the column's type, so a numeric column needs a numeric placeholder.

      Covered in Masking as a guardrail in the data layer

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The mask is a SQL UDF that takes the column value as input and returns the original value or a masked version.”
      ↩︎ Masking as a guardrail in the data layer
      “Policy logic is also evaluated more efficiently than table-specific UDFs.”
      ↩︎ Choosing the layer: ABAC, table-level masks or dynamic views
      “When the query engine must choose between optimization and protecting against information leakage from filtered or masked values, it always makes the secure choice”
      ↩︎ Designing masks that meet a performance objective
      “Expressions that can throw errors (such as ANSI division) prevent the SQL compiler from pushing operations down in the query plan”
      ↩︎ Designing masks that meet a performance objective
      “The mask is a SQL UDF that takes the column value as input and returns the original value or a masked version.”
      ↩︎ Key concept
      “The return type must match or be castable to the column's data type.”
      ↩︎ Exam trap 4
    2. 2.
      “individual table owners can't remove, modify, or bypass it”
      ↩︎ Choosing the layer: ABAC, table-level masks or dynamic views
      “Because they lack a SecureView barrier, they don't protect against probing attacks”
      ↩︎ Exam trap 1
      “Dynamic views fully support query optimization and predicate pushdown, so they can offer better query performance than row filters and column masks.”
      ↩︎ Prediction
      “Because they lack a SecureView barrier, they don't protect against probing attacks”
      ↩︎ Checkpoint
    3. 3.
      “Regex-based masking on serialized documents forces the engine to scan and rewrite the entire payload for every row.”
      ↩︎ Designing masks that meet a performance objective
      “Each distinct mask on a table adds overhead; reusing the same function across columns can reduce it.”
      ↩︎ Exam trap 3
      “These functions are resolved once during query analysis, not per row.”
      ↩︎ Checkpoint
      “The EXCEPT clause eliminates the policy entirely for exempt users, which means no UDF execution for those users.”
      ↩︎ Checkpoint
    4. 4.
      “For tables with column masks, you can exclude masked columns from the index using the columns to sync setting.”
      ↩︎ Where source-table masks stop: indexes, pipelines and requests
      “Exempted principals see unfiltered, unmasked data.”
      ↩︎ Where source-table masks stop: indexes, pipelines and requests
      “ABAC policies on a source table don't apply to AI Search indexes created from that table.”
      ↩︎ Exam trap 2
      “the materialized view or streaming table permanently contains masked or filtered data”
      ↩︎ Checkpoint
    5. 5.
      “Redaction should catch as much sensitive data as possible, since a missed value is worse than an over-redaction, so it is optimized for recall.”
      ↩︎ Where source-table masks stop: indexes, pipelines and requests
      “Detection is deterministic and fast. It adds well under 50 ms to a request, even on a large (100-turn) conversation.”
      ↩︎ Where source-table masks stop: indexes, pipelines and requests

    Spotted a mistake, or was something unclear? Tell us.