What you will be able to do
- Explain why problematic text (harmful language, PII, poisoned documents) is best handled in the data pipeline before it reaches a GenAI application
- Recommend filtering, in-place masking, or record-dropping expectations as alternatives, and justify which one fits a given scenario
- Use ai_classify and ai_mask to produce filter signals and masked text from SQL or PySpark
- Identify when an inference-time redaction policy is a fallback, and what it cannot detect
Key concept
Pipeline-stage filtering of problematic text — Harmful, sensitive or poisoned content in a source corpus is handled as a deliberate preprocessing step. A classifier or detector scores each document, and that score decides whether the document is removed, rewritten or kept before chunking and embedding.
1.Where problematic text enters a GenAI application
A retrieval-augmented application can only answer from what its corpus contains. So whatever is in the corpus can come back out: an abusive forum post, a support ticket with a customer's email address, a planted document meant to mislead the model. The Databricks RAG data-pipeline guidance names three kinds of content to worry about. Some documents are irrelevant or out of date. Some contain problematic content such as harmful language. Others hold sensitive information that you do not want the agent to expose. On top of these, every source is a possible attack surface: any document you feed in can be used by bad actors for data poisoning.
The guidance puts the answer in the pipeline itself. A RAG pipeline runs in stages: ingest, preprocess (parse, enrich, deduplicate, filter), chunk, embed, index. Filtering is listed as a preprocessing step with a specific purpose: removing irrelevant or unwanted documents from the collection. Problematic-text mitigation belongs there, after parsing and before chunking. Nothing removed at that stage is ever embedded, indexed or retrieved.
Two habits from the same guidance make that step safe to run. First, ingest raw source data into a target table, and filter downstream of it. You keep the original for preservation, traceability and auditing, so a filtering mistake can be corrected without re-ingesting. Second, clean the parsed text so the chain is not fed noise such as headers, footers or stray special characters. That is a different job from removing harmful content, but it happens in the same preprocessing stage.
Checkpoint 1 of 7· Check yourself
Why does the Databricks RAG pipeline guidance recommend ingesting raw source data into a target table before filtering it?
Storing the raw data in a target table keeps the original for preservation, traceability and auditing. A filtering mistake can then be fixed downstream without going back to the source.
“raw source data should be ingested and stored in a target table. This approach ensures data preservation, traceability, and auditing.”Source: docs.databricks.com
You also need to know where the sensitive material is before choosing a mitigation. For tables governed in Unity Catalog, Databricks Data Classification uses an agent to classify and tag tables automatically. Those results tell you which source tables carry sensitive classes, and scanning is incremental, so new tables are typically scanned within 24 hours of creation. Classification tells you what is there. It does not change the text. Changing the text is the job of the alternatives in the rest of this lesson.
2.Alternative 1: Filter whole documents out with a classifier
The most direct mitigation is to remove a document from the corpus. The guidance says to add a pipeline step that filters these documents using any metadata. Its first example is a toxicity classifier whose prediction becomes the filter condition. Its second is a PII detection algorithm run over the documents to filter them. The same reasoning applies to poisoning: the guidance suggests adding detection and filtering mechanisms to find and remove poisoned documents.
On Databricks, one way to produce that prediction is an AI function. ai_classify assigns each document to one of the labels you supply, using an LLM. It exists as a PySpark function and as a corresponding Databricks SQL function. The labels can be a fixed list or a dictionary that maps each label to a description, so you define categories that fit your content policy. You then keep only the rows whose label is acceptable.
# Static labels (same set for every row)
df.select(ai_classify("text", ["positive", "negative", "neutral"]))Store the classifier's output as metadata on the document. The guidance treats content-based tags, including categories like PII or HIPAA, as normal enrichment. Once the label is stored as a column, the filter is a simple predicate, and you can audit or re-run it later without classifying the corpus again.
Filtering has one clear cost: the whole document is gone. That is the right result for content with no value to the application, such as abusive text, irrelevant material or suspected poisoning. It is a heavy-handed result for a useful document that happens to contain a phone number.
Checkpoint 2 of 7· Check yourself
A team building a customer-support RAG agent finds that some scraped community posts contain abusive language. What does the Databricks RAG pipeline guidance recommend?
The guidance makes filtering an explicit preprocessing step that removes unwanted documents. Its example is a toxicity classifier whose prediction serves as the filter. Chunking and deduplication solve other problems.
“applying a toxicity classifier to the document to produce a prediction you can use as a filter”Source: docs.databricks.com
3.Alternative 2: Mask sensitive entities and keep the document
When the problem is a few sensitive values inside otherwise useful text, rewrite the text instead of removing it. ai_mask(content, labels) calls an AI model through a Foundation Model APIs chat endpoint to mask the entity types you list. content is a STRING. labels is an ARRAY<STRING> literal in which each element names a type of information to mask. The function returns the string with those entities replaced. NULL input returns NULL.
> SELECT ai_mask(
'John Doe lives in New York. His email is john.doe@example.com.',
array('person', 'email')
);
"[MASKED] lives in New York. His email is [MASKED]."In the example, "New York" survives because 'location' was not in the label array. You decide exactly what gets masked. Because ai_mask uses a model, it can find entities such as a person's name, which have no fixed format.
The constraints matter when you recommend it. The function is in Public Preview. Although the underlying model can handle several languages, it is tuned for English. It requires Databricks Runtime 15.4 LTS or above, it is not available on Databricks SQL Classic, and it runs only in regions that support AI Functions optimized for batch inference. A multilingual corpus or a Classic-only SQL environment can therefore rule it out.
Checkpoint 3 of 7· Check yourself
You run ai_mask over a document column with array('person', 'email'). Which statement about the result is correct?
ai_mask returns the same text with only the requested entity types masked. It removes no rows and does not return a flag. In the documented example, the location 'New York' is left unmasked.
“A STRING where the specified information is masked.”Source: docs.databricks.com
Checkpoint 4 of 7· Exam question
An internal HR assistant retrieves context from a Delta table built from years of archived employee emails. A review finds that many emails contain employee social security numbers and salary figures embedded in the body text. What is the best way to mitigate this before the data feeds the RAG pipeline?
Correct answer: A — Add a PII detection and redaction step to the data preparation pipeline so sensitive values are masked or removed before the text is embedded and indexed
- A. Detecting and redacting PII during data preparation removes sensitive values from the source text before it is ever embedded, preventing SSNs and salaries from entering the retrievable index at all.
- B. An output guardrail only inspects the model's final response; the sensitive values remain embedded in the vector index and can still be retrieved into the model's context window during generation.
- C. Restricting access to the index controls who can query it, but it does not remove the SSNs and salaries from the archived email text, so the underlying problematic content is still present.
- D. Encrypting the table at rest protects against unauthorized storage access but does not strip PII from the text, so the sensitive values are still exposed to anyone permitted to query the index.
Sources4
4.Alternative 3: Enforce the filter as a pipeline expectation
If the corpus is built with Lakeflow pipelines as streaming tables or materialized views, the filter can be declared as a data-quality rule instead of written as ad hoc code. Expectations are SQL Boolean constraints applied to every record. They let you fail an update or drop records when invalid ones are detected.
The operator you choose matters. The default expect keeps violating records in the target dataset and only records metrics on how many passed or failed. That is useful for measuring how much problematic text you have, but it removes nothing. expect_or_drop stops further processing of violating records: they are dropped from the target dataset. You can still see the drop counts on the pipeline's Data quality tab or in the event log, which gives you a record of how much was filtered.
CONSTRAINT valid_current_page EXPECT (current_page_id IS NOT NULL and current_page_title IS NOT NULL) ON VIOLATION DROP ROWA constraint must be valid SQL. It cannot contain custom Python functions, external service calls, or subqueries that reference other tables. So the expectation cannot run your toxicity or PII classifier itself. In practice, the classifier's prediction is written to a column at an earlier stage (Alternative 1). The expectation then tests that column. The classifier produces the label, and the expectation enforces it and counts what it drops.
Checkpoint 5 of 7· Fill the gap
Which operator makes this expectation remove violating records from the target dataset, instead of keeping them and only recording metrics?
@dp. ? ("valid_current_page", "current_page_id IS NOT NULL AND current_page_title IS NOT NULL")expect_or_drop drops records that violate the constraint. Plain expect keeps them in the target dataset and only tracks pass/fail metrics.
Source: docs.databricks.comSources5
5.When the source can't be cleaned: redaction at the gateway, and choosing between alternatives
Sometimes you cannot rewrite the source. The data may belong to another team, or it may arrive faster than a batch job can process it. In that case a backstop sits in front of the model. Sensitive Data Detection is a built-in service policy (in Beta) that you attach to a model or model provider service in Unity Gateway. It inspects requests and responses and either blocks the interaction or redacts the matched value in place. A redacted value becomes a category placeholder such as [US_SSN]. Compare that with ai_mask, which writes a generic [MASKED].
The policy is deterministic. It uses regex, checksums and nearby context keywords across 15 structured categories such as email, credit card and US SSN, and it adds well under 50 ms per request. Because it relies on patterns, it cannot detect names, locations, organizations or other free-text entities. For those, Databricks recommends adding an LLM-judge guardrail at a higher rank, so that it evaluates content the detector has already redacted. This is a runtime safety net. It does not clean the source: the corpus still holds the sensitive text, and harmful language or poisoned documents are outside what this policy detects.
| Alternative | What happens to the text | Where it runs | Recommend when |
|---|---|---|---|
| Classifier filter (e.g. ai_classify toxicity/PII label) | Whole document removed from the corpus | Preprocessing, before chunking | Content has no value to the app: harmful language, irrelevant or suspected poisoned documents |
| ai_mask | Named entity types replaced with [MASKED]; document kept | Preprocessing via SQL | A useful document contains a few sensitive values, including names |
| expect_or_drop expectation | Violating records dropped, with metrics recorded | Lakeflow pipeline table definition | The pipeline must enforce an existing label column and you need drop counts |
| Sensitive Data Detection policy | Matched value blocked or redacted, e.g. [US_SSN] | Unity Gateway, on requests/responses | The source can't be changed; only structured PII patterns need to be caught |
Checkpoint 6 of 7· Match them up
Match each mitigation to how it treats problematic text
Tap a term, then the definition that fits it.
Filtering and expectations both remove data, at document and record level respectively. ai_mask rewrites the source text. Sensitive Data Detection acts only on traffic through the gateway, at inference time.
“Databricks replaces each matched value with a placeholder token, such as [US_SSN], and forwards the rewritten content.”Source: docs.databricks.com
Checkpoint 7 of 7· Exam question
A product team is building a RAG assistant using articles scraped from third-party blogs with unclear reuse terms. Legal flags that the scraped corpus creates copyright and licensing risk for the company. What should the team recommend to mitigate this risk in the data source?
Correct answer: A — Replace the scraped blog corpus with first-party content or a dataset under a license that explicitly permits this use
- A. Swapping in first-party content or a dataset with a license that clearly permits the intended use removes the legal risk at its source, since the corpus no longer relies on ambiguous reuse terms.
- B. Citing the source URL may be good practice for attribution, but it does not resolve the underlying licensing uncertainty about whether the company is permitted to reuse and redistribute the scraped content.
- C. Shortening the chunk size changes how much text is retrieved per query, but the content still originates from a corpus with unresolved reuse rights, so the licensing risk remains unchanged.
- D. A fair-use disclaimer in the prompt is not a legal determination and does not change the actual licensing status of the scraped content, so the underlying risk to the company persists.
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Problematic text in a source corpus is best handled only at inference time, with a guardrail on the model's output.Why is that wrong?
The RAG pipeline guidance puts the mitigation in the data pipeline: add a filtering step so unwanted documents are never chunked, embedded or retrieved. Inference-time policies are a backstop.
Covered in Where problematic text enters a GenAI application
2.If a document contains PII, the only safe option is to drop the whole document.Why is that wrong?
ai_mask replaces only the entity types you specify and returns the rest of the text. A useful document can be kept with its sensitive values masked.
Covered in Alternative 2: Mask sensitive entities and keep the document
3.A plain expect constraint removes the records that violate it.Why is that wrong?
Plain expect keeps violating records and only records metrics. Records are dropped only with expect_or_drop (ON VIOLATION DROP ROW). A constraint also cannot call external services, so the classifier must run in an earlier step.
Covered in Alternative 3: Enforce the filter as a pipeline expectation
4.Gateway Sensitive Data Detection will also redact people's names found in retrieved documents.Why is that wrong?
The policy matches structured formats with regex, checksums and context keywords. Names and other free-text entities need a model, such as an LLM-judge guardrail or ai_mask.
Covered in When the source can't be cleaned: redaction at the gateway, and choosing between alternatives
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“they contain problematic content such as harmful language”
↩︎ Where problematic text enters a GenAI application“any document sources you feed into your agent are potential attack vectors for bad actors to launch data poisoning attacks”
↩︎ Where problematic text enters a GenAI application“raw source data should be ingested and stored in a target table. This approach ensures data preservation, traceability, and auditing.”
↩︎ Where problematic text enters a GenAI application“applying a personally identifiable information (PII) detection algorithm to the documents to filter documents”
↩︎ Alternative 1: Filter whole documents out with a classifier“applying a toxicity classifier to the document to produce a prediction you can use as a filter”
↩︎ Key concept“consider including a step in your pipeline to filter out these documents by using any metadata”
↩︎ Exam trap 1“applying a toxicity classifier to the document to produce a prediction you can use as a filter”
↩︎ Checkpoint - 2.
“Databricks Data Classification uses an agent to automatically classify and tag tables in your catalog.”
↩︎ Where problematic text enters a GenAI application - 3.
“Classifies document content into one of the provided labels using AI/LLM.”
↩︎ Alternative 1: Filter whole documents out with a classifier - 4.
“this AI Function is tuned for English”
↩︎ Alternative 2: Mask sensitive entities and keep the document“The ai_mask() function allows you to invoke a state-of-the-art AI model to mask specified entities in a given text using SQL.”
↩︎ Exam trap 2“A STRING where the specified information is masked.”
↩︎ Prediction - 5.https://docs.databricks.com/aws/en/ldp/expectationsOfficial docs
“Use the expect_or_drop operator to prevent further processing of invalid records.”
↩︎ Alternative 3: Enforce the filter as a pipeline expectation“Constraints must use valid SQL syntax and cannot contain the following:”
↩︎ Alternative 3: Enforce the filter as a pipeline expectation“Records that violate the expectation are added to the target dataset along with valid records”
↩︎ Exam trap 3 - 6.https://docs.databricks.com/aws/en/data-governance/unity-catalog/service-policies/detect-sensitive-dataOfficial docs
“either blocks the interaction or redacts the matched value in place”
↩︎ When the source can't be cleaned: redaction at the gateway, and choosing between alternatives“Keep Sensitive Data Detection at a lower rank to redact structured data.”
↩︎ When the source can't be cleaned: redaction at the gateway, and choosing between alternatives“Names, locations, and organizations, and other free-text entities that need language understanding to identify.”
↩︎ Exam trap 4“Databricks replaces each matched value with a placeholder token, such as [US_SSN], and forwards the rewritten content.”
↩︎ Checkpoint