CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 5 · Lesson 46/56

    Mitigating Problematic Text in GenAI Data Sources

    Recommend an alternative for problematic text mitigation in a data source feeding a GenAI application

    15 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why problematic text (harmful language, PII, poisoned documents) is best handled in the data pipeline before it reaches a GenAI application
    • Recommend filtering, in-place masking, or record-dropping expectations as alternatives, and justify which one fits a given scenario
    • Use ai_classify and ai_mask to produce filter signals and masked text from SQL or PySpark
    • Identify when an inference-time redaction policy is a fallback, and what it cannot detect

    Key concept

    Pipeline-stage filtering of problematic text — Harmful, sensitive or poisoned content in a source corpus is handled as a deliberate preprocessing step. A classifier or detector scores each document, and that score decides whether the document is removed, rewritten or kept before chunking and embedding.

    1.Where problematic text enters a GenAI application

    A retrieval-augmented application can only answer from what its corpus contains. So whatever is in the corpus can come back out: an abusive forum post, a support ticket with a customer's email address, a planted document meant to mislead the model. The Databricks RAG data-pipeline guidance names three kinds of content to worry about. Some documents are irrelevant or out of date. Some contain problematic content such as harmful language. Others hold sensitive information that you do not want the agent to expose. On top of these, every source is a possible attack surface: any document you feed in can be used by bad actors for data poisoning.

    The guidance puts the answer in the pipeline itself. A RAG pipeline runs in stages: ingest, preprocess (parse, enrich, deduplicate, filter), chunk, embed, index. Filtering is listed as a preprocessing step with a specific purpose: removing irrelevant or unwanted documents from the collection. Problematic-text mitigation belongs there, after parsing and before chunking. Nothing removed at that stage is ever embedded, indexed or retrieved.

    Two habits from the same guidance make that step safe to run. First, ingest raw source data into a target table, and filter downstream of it. You keep the original for preservation, traceability and auditing, so a filtering mistake can be corrected without re-ingesting. Second, clean the parsed text so the chain is not fed noise such as headers, footers or stray special characters. That is a different job from removing harmful content, but it happens in the same preprocessing stage.

    Checkpoint 1 of 7· Check yourself

    Why does the Databricks RAG pipeline guidance recommend ingesting raw source data into a target table before filtering it?

    You also need to know where the sensitive material is before choosing a mitigation. For tables governed in Unity Catalog, Databricks Data Classification uses an agent to classify and tag tables automatically. Those results tell you which source tables carry sensitive classes, and scanning is incremental, so new tables are typically scanned within 24 hours of creation. Classification tells you what is there. It does not change the text. Changing the text is the job of the alternatives in the rest of this lesson.

    Sources12

    2.Alternative 1: Filter whole documents out with a classifier

    The most direct mitigation is to remove a document from the corpus. The guidance says to add a pipeline step that filters these documents using any metadata. Its first example is a toxicity classifier whose prediction becomes the filter condition. Its second is a PII detection algorithm run over the documents to filter them. The same reasoning applies to poisoning: the guidance suggests adding detection and filtering mechanisms to find and remove poisoned documents.

    On Databricks, one way to produce that prediction is an AI function. ai_classify assigns each document to one of the labels you supply, using an LLM. It exists as a PySpark function and as a corresponding Databricks SQL function. The labels can be a fixed list or a dictionary that maps each label to a description, so you define categories that fit your content policy. You then keep only the rows whose label is acceptable.

    ai_classify with a static label set. The function returns one label per row, and that label can be stored as a column and used as a filter.python
    # Static labels (same set for every row)
    df.select(ai_classify("text", ["positive", "negative", "neutral"]))

    Store the classifier's output as metadata on the document. The guidance treats content-based tags, including categories like PII or HIPAA, as normal enrichment. Once the label is stored as a column, the filter is a simple predicate, and you can audit or re-run it later without classifying the corpus again.

    Filtering has one clear cost: the whole document is gone. That is the right result for content with no value to the application, such as abusive text, irrelevant material or suspected poisoning. It is a heavy-handed result for a useful document that happens to contain a phone number.

    Checkpoint 2 of 7· Check yourself

    A team building a customer-support RAG agent finds that some scraped community posts contain abusive language. What does the Databricks RAG pipeline guidance recommend?

    Sources13

    3.Alternative 2: Mask sensitive entities and keep the document

    When the problem is a few sensitive values inside otherwise useful text, rewrite the text instead of removing it. ai_mask(content, labels) calls an AI model through a Foundation Model APIs chat endpoint to mask the entity types you list. content is a STRING. labels is an ARRAY<STRING> literal in which each element names a type of information to mask. The function returns the string with those entities replaced. NULL input returns NULL.

    ai_mask replaces the requested entity types with [MASKED] and leaves the rest of the sentence intact.sql
    > SELECT ai_mask(
        'John Doe lives in New York. His email is john.doe@example.com.',
        array('person', 'email')
      );
     "[MASKED] lives in New York. His email is [MASKED]."

    In the example, "New York" survives because 'location' was not in the label array. You decide exactly what gets masked. Because ai_mask uses a model, it can find entities such as a person's name, which have no fixed format.

    The constraints matter when you recommend it. The function is in Public Preview. Although the underlying model can handle several languages, it is tuned for English. It requires Databricks Runtime 15.4 LTS or above, it is not available on Databricks SQL Classic, and it runs only in regions that support AI Functions optimized for batch inference. A multilingual corpus or a Classic-only SQL environment can therefore rule it out.

    Checkpoint 3 of 7· Check yourself

    You run ai_mask over a document column with array('person', 'email'). Which statement about the result is correct?

    Checkpoint 4 of 7· Exam question

    An internal HR assistant retrieves context from a Delta table built from years of archived employee emails. A review finds that many emails contain employee social security numbers and salary figures embedded in the body text. What is the best way to mitigate this before the data feeds the RAG pipeline?

    Sources4

    4.Alternative 3: Enforce the filter as a pipeline expectation

    If the corpus is built with Lakeflow pipelines as streaming tables or materialized views, the filter can be declared as a data-quality rule instead of written as ad hoc code. Expectations are SQL Boolean constraints applied to every record. They let you fail an update or drop records when invalid ones are detected.

    The operator you choose matters. The default expect keeps violating records in the target dataset and only records metrics on how many passed or failed. That is useful for measuring how much problematic text you have, but it removes nothing. expect_or_drop stops further processing of violating records: they are dropped from the target dataset. You can still see the drop counts on the pipeline's Data quality tab or in the event log, which gives you a record of how much was filtered.

    SQL form of a dropping expectation: rows that fail the constraint never reach the target table.sql
    CONSTRAINT valid_current_page EXPECT (current_page_id IS NOT NULL and current_page_title IS NOT NULL) ON VIOLATION DROP ROW

    A constraint must be valid SQL. It cannot contain custom Python functions, external service calls, or subqueries that reference other tables. So the expectation cannot run your toxicity or PII classifier itself. In practice, the classifier's prediction is written to a column at an earlier stage (Alternative 1). The expectation then tests that column. The classifier produces the label, and the expectation enforces it and counts what it drops.

    Checkpoint 5 of 7· Fill the gap

    Which operator makes this expectation remove violating records from the target dataset, instead of keeping them and only recording metrics?

    @dp. ? ("valid_current_page", "current_page_id IS NOT NULL AND current_page_title IS NOT NULL")

    Sources5

    5.When the source can't be cleaned: redaction at the gateway, and choosing between alternatives

    Sometimes you cannot rewrite the source. The data may belong to another team, or it may arrive faster than a batch job can process it. In that case a backstop sits in front of the model. Sensitive Data Detection is a built-in service policy (in Beta) that you attach to a model or model provider service in Unity Gateway. It inspects requests and responses and either blocks the interaction or redacts the matched value in place. A redacted value becomes a category placeholder such as [US_SSN]. Compare that with ai_mask, which writes a generic [MASKED].

    The policy is deterministic. It uses regex, checksums and nearby context keywords across 15 structured categories such as email, credit card and US SSN, and it adds well under 50 ms per request. Because it relies on patterns, it cannot detect names, locations, organizations or other free-text entities. For those, Databricks recommends adding an LLM-judge guardrail at a higher rank, so that it evaluates content the detector has already redacted. This is a runtime safety net. It does not clean the source: the corpus still holds the sensitive text, and harmful language or poisoned documents are outside what this policy detects.

    The four alternatives side by side
    AlternativeWhat happens to the textWhere it runsRecommend when
    Classifier filter (e.g. ai_classify toxicity/PII label)Whole document removed from the corpusPreprocessing, before chunkingContent has no value to the app: harmful language, irrelevant or suspected poisoned documents
    ai_maskNamed entity types replaced with [MASKED]; document keptPreprocessing via SQLA useful document contains a few sensitive values, including names
    expect_or_drop expectationViolating records dropped, with metrics recordedLakeflow pipeline table definitionThe pipeline must enforce an existing label column and you need drop counts
    Sensitive Data Detection policyMatched value blocked or redacted, e.g. [US_SSN]Unity Gateway, on requests/responsesThe source can't be changed; only structured PII patterns need to be caught

    Checkpoint 6 of 7· Match them up

    Match each mitigation to how it treats problematic text

    Tap a term, then the definition that fits it.

    Checkpoint 7 of 7· Exam question

    A product team is building a RAG assistant using articles scraped from third-party blogs with unclear reuse terms. Legal flags that the scraped corpus creates copyright and licensing risk for the company. What should the team recommend to mitigate this risk in the data source?

    Sources6

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Problematic text in a source corpus is best handled only at inference time, with a guardrail on the model's output.Why is that wrong?

      The RAG pipeline guidance puts the mitigation in the data pipeline: add a filtering step so unwanted documents are never chunked, embedded or retrieved. Inference-time policies are a backstop.

      Covered in Where problematic text enters a GenAI application

    2. 2.If a document contains PII, the only safe option is to drop the whole document.Why is that wrong?

      ai_mask replaces only the entity types you specify and returns the rest of the text. A useful document can be kept with its sensitive values masked.

      Covered in Alternative 2: Mask sensitive entities and keep the document

    3. 3.A plain expect constraint removes the records that violate it.Why is that wrong?

      Plain expect keeps violating records and only records metrics. Records are dropped only with expect_or_drop (ON VIOLATION DROP ROW). A constraint also cannot call external services, so the classifier must run in an earlier step.

      Covered in Alternative 3: Enforce the filter as a pipeline expectation

    4. 4.Gateway Sensitive Data Detection will also redact people's names found in retrieved documents.Why is that wrong?

      The policy matches structured formats with regex, checksums and context keywords. Names and other free-text entities need a model, such as an LLM-judge guardrail or ai_mask.

      Covered in When the source can't be cleaned: redaction at the gateway, and choosing between alternatives

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “they contain problematic content such as harmful language”
      ↩︎ Where problematic text enters a GenAI application
      “any document sources you feed into your agent are potential attack vectors for bad actors to launch data poisoning attacks”
      ↩︎ Where problematic text enters a GenAI application
      “raw source data should be ingested and stored in a target table. This approach ensures data preservation, traceability, and auditing.”
      ↩︎ Where problematic text enters a GenAI application
      “applying a personally identifiable information (PII) detection algorithm to the documents to filter documents”
      ↩︎ Alternative 1: Filter whole documents out with a classifier
      “applying a toxicity classifier to the document to produce a prediction you can use as a filter”
      ↩︎ Key concept
      “consider including a step in your pipeline to filter out these documents by using any metadata”
      ↩︎ Exam trap 1
      “applying a toxicity classifier to the document to produce a prediction you can use as a filter”
      ↩︎ Checkpoint
    2. 2.
      “Databricks Data Classification uses an agent to automatically classify and tag tables in your catalog.”
      ↩︎ Where problematic text enters a GenAI application
    3. 3.
    4. 4.
      “The ai_mask() function allows you to invoke a state-of-the-art AI model to mask specified entities in a given text using SQL.”
      ↩︎ Exam trap 2
      “A STRING where the specified information is masked.”
      ↩︎ Prediction
    5. 5.
      “Use the expect_or_drop operator to prevent further processing of invalid records.”
      ↩︎ Alternative 3: Enforce the filter as a pipeline expectation
      “Constraints must use valid SQL syntax and cannot contain the following:”
      ↩︎ Alternative 3: Enforce the filter as a pipeline expectation
      “Records that violate the expectation are added to the target dataset along with valid records”
      ↩︎ Exam trap 3
    6. 6.
      “either blocks the interaction or redacts the matched value in place”
      ↩︎ When the source can't be cleaned: redaction at the gateway, and choosing between alternatives
      “Keep Sensitive Data Detection at a lower rank to redact structured data.”
      ↩︎ When the source can't be cleaned: redaction at the gateway, and choosing between alternatives
      “Names, locations, and organizations, and other free-text entities that need language understanding to identify.”
      ↩︎ Exam trap 4
      “Databricks replaces each matched value with a placeholder token, such as [US_SSN], and forwards the rewritten content.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.