CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 8/56

    Filtering Extraneous Content from RAG Source Documents

    Filter extraneous content in source documents that degrades quality of a RAG application

    15 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain how boilerplate, noise, duplicates and unwanted documents degrade RAG retrieval and generation quality
    • Remove headers, footers, page numbers and special characters during parsing while keeping useful document structure
    • Use ai_parse_document element types to separate layout markers from body content
    • Filter whole documents that are irrelevant, outdated, toxic, sensitive or possibly poisoned
    • Describe the MinHash process for removing duplicate and near-duplicate documents before chunking

    Key concept

    Superfluous context — Any text in the corpus that does not help answer a query, such as repeated headers, footers, page numbers, duplicate copies or off-topic documents. It gets retrieved alongside real content, takes up space in the LLM's context, and lowers both retrieval and generation quality. That is why it is removed in the data pipeline, before chunking.

    1.Why extraneous content is a retrieval problem first

    Databricks sorts RAG quality problems into two kinds. Retrieval quality asks whether you fetched the most relevant information for a query. Generation quality asks whether the LLM wrote an accurate, helpful answer from what it was given. Extraneous content harms both, but the harm starts at retrieval. Text that is irrelevant, malformed or repeated becomes part of the chunks, then gets embedded and returned as context, and the model has to work around it.

    The Databricks unstructured data pipeline puts the noise-removal steps under data preprocessing, ahead of chunking, embedding and indexing. Four of those preprocessing steps are about getting rid of content you don't want:

    - Parsing: extract the relevant information from raw files and clean it. - Enrichment: add metadata *and remove noise*. - Deduplication: find and remove duplicate or near-duplicate documents. - Filtering: remove irrelevant or unwanted documents from the collection.

    The order matters. The pipeline guide says chunking happens only after the data has been parsed, deduplicated and filtered. Chunking also has its own role here: splitting documents into semantically focused chunks helps keep distracting text out of what gets retrieved. Any noise still present at that stage ends up embedded in every chunk it touches.

    Checkpoint 1 of 5· Check yourself

    According to the Databricks unstructured data pipeline, when should duplicate removal and document filtering happen relative to chunking?

    Sources12

    2.Cleaning boilerplate and noise during parsing

    Parsing is the first and cheapest point to remove noise, because this is where you decide which parts of the raw file become text. The pipeline guide names data cleaning as a parsing best practice: remove headers, footers and special characters so the RAG chain has less unnecessary or malformed text to process. The AI Search retrieval quality guide gives a short cleaning checklist with three items:

    1. Remove boilerplate (headers, footers, page numbers). 2. Preserve document structure (headings, lists, tables). 3. Maintain semantic boundaries when chunking.

    Item 2 limits item 1. Cleaning removes repeated page furniture, not the structure that carries meaning. Headings, lists and tables stay, because they help with chunking later and can become structural metadata.

    Which tool does the cleaning depends on the source format. For PDFs and Word documents, the pipeline guide points to off-the-shelf libraries such as unstructured and PyPDF2, which offer options for customizing the parsing. For HTML, it points to BeautifulSoup and lxml. These let you navigate the page structure and select only the elements you want, which is how you leave out everything around the main content. Scanned documents and images need OCR (Optical Character Recognition) first. For PDFs and complex documents, Databricks recommends ai_parse_document for clean, structured text. The retrieval quality guide also warns that poor parsing, such as missing tables or broken formatting, directly lowers retrieval quality.

    Kinds of extraneous content and the pipeline step that removes each one
    Extraneous contentWhere it is removedMechanism in the sources
    Headers, footers, page numbersParsing (data cleaning)Strip boilerplate; with ai_parse_document, use the page_header, page_footer and page_number element types
    Special characters, malformed textParsing (data cleaning)Preprocess extracted text so the chain processes less malformed information
    Content around the main body of a web pageParsing (HTML)BeautifulSoup or lxml to select specific elements
    Duplicate or near-duplicate documentsDeduplicationMetadata matching, then MinHash locality-sensitive hashing
    Irrelevant, outdated, toxic or PII-bearing documentsFilteringMetadata, a toxicity classifier, or a PII detection algorithm

    Two more parsing practices help catch noise you didn't expect. Customize the parsing logic to fit your documents' structure; the extra work up front is worth it because it prevents downstream quality problems. Evaluate parsing quality on a regular basis by manually reviewing a sample of the parsed output. That is often how you find a repeated banner or a garbled table that an automated step missed. Error handling and logging during parsing also help, and they often reveal upstream problems in the source data.

    Checkpoint 2 of 5· Check yourself

    A team cleans PDFs before chunking by removing every heading, list marker and table so that only plain sentences remain. Which Databricks guidance does this break?

    Sources23

    3.Separating layout markers from body text with ai_parse_document

    ai_parse_document makes boilerplate removal easier because it labels each piece of a document. It identifies layout information such as page numbers, headers, tables and footers, and returns the document as an array of elements. Each element is one distinct block of content: a paragraph, a table, a figure, or a layout marker such as a page header or footer. Every element carries a type field, so you can filter on that field instead of writing regular expressions to guess which lines are footers.

    Part of the ai_parse_document output schema (v2.0): each element has a type, its content, and a confidence scorejson
        "elements": [
          {
            "id": INT,                 // 0-based element index
            "type": STRING,            // Supported: text, table, figure, table, title, caption, section_header,
                                       // page_footer, page_header, page_number, footnote
            "content": STRING,         // Text content of the target element
            "confidence": DOUBLE,      // Confidence score of the target element

    Three element types match the boilerplate in the cleaning checklist: page_header, page_footer and page_number. These are the obvious ones to drop before chunking. section_header is a different thing. It marks the start of a section, which is the kind of structure the checklist says to keep, and the retrieval quality guide suggests adding section headers to chunks as semantic metadata. Table elements come back as HTML, so table contents are kept instead of being flattened into broken text. Each element also has a confidence score for how reliably it was extracted, which is a useful signal when you review parsing quality. The function can also be limited to certain pages with pageRange, and it fails on documents over 500 pages unless a range is given.

    Checkpoint 3 of 5· Match them up

    Match each ai_parse_document element type to what it represents

    Tap a term, then the definition that fits it.

    Sources43

    4.Filtering out whole documents that don't belong

    Some extraneous content is an entire document, not a few lines inside one. The filtering step removes irrelevant or unwanted documents from the collection. The pipeline guide lists the reasons a document may not be useful to your agent: it is irrelevant to the agent's purpose, it is too old or unreliable, or it contains problematic content such as harmful language. Other documents may contain sensitive information that you don't want the agent to expose.

    The guide's answer is to filter on metadata, including metadata that a model produces. For example, run a toxicity classifier on each document and use its prediction as a filter, or run a PII (personally identifiable information) detection algorithm and filter documents based on what it finds. Ordinary metadata such as dates and source system can handle the "too old" and "unreliable" cases.

    There is also a security reason. Every document source you feed into the agent is a possible attack vector for data poisoning, so the guide suggests adding detection and filtering mechanisms to find and remove poisoned documents. Filtering protects safety and trust as well as answer quality.

    Checkpoint 4 of 5· Check yourself

    A pipeline ingests shared-drive documents for an HR assistant. Some contain employee personal data. What does Databricks suggest adding?

    Sources2

    5.Removing duplicate and near-duplicate documents

    Duplicates are a quieter kind of noise. If you pull from several shared drives, the same document may exist in several places, sometimes with small edits. A knowledge base may also hold copies of product documentation or drafts of blog posts. If the duplicates stay in, the index fills with highly redundant chunks, which lowers application performance. A query can then return the same passage several times, crowding out other relevant passages.

    Start with metadata. If items share a title and creation date but come from different sources or locations, you can filter them on metadata alone. That isn't enough for copies with small edits. For those, the guide recommends locality-sensitive hashing, specifically MinHash, which has a Spark implementation in Spark ML. MinHash builds a hash from the words a document contains, then finds duplicates and near-duplicates by joining on those hashes.

    Checkpoint 5 of 5· Put it in order

    Put the four steps of MinHash deduplication in order

    1. 1.Run a similarity join on the hashes to find each duplicate or near-duplicate
    2. 2.Create a feature vector for each document (optionally remove stop words, stem or lemmatize, then tokenize into n-grams)
    3. 3.Filter out the duplicates you don't want to keep
    4. 4.Fit a MinHash model and hash the vectors using MinHash for Jaccard distance

    Once duplicates are grouped, you have to decide which copy to keep. A baseline approach picks one arbitrarily, such as the first result or a random choice. A better approach keeps the best version, for example the most recently updated, the published one, or the one from the most authoritative source. Expect to tune the featurization step and the number of hash tables in the MinHash model to improve matching.

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Noisy source text is only a generation problem, so the fix is a better prompt or a stronger LLM.Why is that wrong?

      Noise shapes what gets retrieved. Databricks names the data pipeline's parsing and chunking strategy as an influence on retrieval quality, so the fix belongs in the pipeline.

      Covered in Why extraneous content is a retrieval problem first

    2. 2.When cleaning ai_parse_document output, drop every header-like element, including section_header.Why is that wrong?

      Only page-level layout markers (page_header, page_footer, page_number) are boilerplate. section_header marks the start of a section, which is structure the guidance says to preserve and add to chunks.

      Covered in Separating layout markers from body text with ai_parse_document

    3. 3.Matching on title and creation date is enough to deduplicate a corpus.Why is that wrong?

      Metadata only catches some duplicates. Near-duplicates with small edits need content-based detection such as MinHash locality-sensitive hashing.

      Covered in Removing duplicate and near-duplicate documents

    4. 4.Filtering extraneous content only means trimming lines inside documents; every ingested document should still be indexed.Why is that wrong?

      The pipeline has a separate filtering step that removes entire documents that are irrelevant, outdated, unreliable, harmful or sensitive.

      Covered in Filtering out whole documents that don't belong

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “It's difficult to generate high-quality RAG output if the context provided to the LLM is missing important information or contains superfluous information.”
      ↩︎ Why extraneous content is a retrieval problem first
      “It's difficult to generate high-quality RAG output if the context provided to the LLM is missing important information or contains superfluous information.”
      ↩︎ Key concept
      “Retrieval quality can be influenced by both the data pipeline (for example, parsing/chunking strategy, metadata strategy, embedding model)”
      ↩︎ Exam trap 1
      “Retrieval quality can be influenced by both the data pipeline (for example, parsing/chunking strategy, metadata strategy, embedding model) and the RAG chain”
      ↩︎ Prediction
    2. 2.
      “Enrichment: Enrich data with additional metadata and remove noise.”
      ↩︎ Why extraneous content is a retrieval problem first
      “minimizing the inclusion of distracting or irrelevant information”
      ↩︎ Why extraneous content is a retrieval problem first
      “Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters.”
      ↩︎ Cleaning boilerplate and noise during parsing
      “Reduce the amount of unnecessary or malformed information your RAG chain needs to process.”
      ↩︎ Cleaning boilerplate and noise during parsing
      “These libraries can help navigate the HTML structure, select specific elements, and extract the desired text or attributes.”
      ↩︎ Cleaning boilerplate and noise during parsing
      “Regularly assess the quality of the parsed data by manually reviewing a sample of the output.”
      ↩︎ Cleaning boilerplate and noise during parsing
      “either because they are irrelevant to its purpose, too old or unreliable, or because they contain problematic content such as harmful language.”
      ↩︎ Filtering out whole documents that don't belong
      “applying a toxicity classifier to the document to produce a prediction you can use as a filter.”
      ↩︎ Filtering out whole documents that don't belong
      “any document sources you feed into your agent are potential attack vectors for bad actors to launch data poisoning attacks.”
      ↩︎ Filtering out whole documents that don't belong
      “you can end up with highly redundant chunks in your final index that can decrease the performance of your application.”
      ↩︎ Removing duplicate and near-duplicate documents
      “Specifically, a technique called MinHash works well here, and a Spark implementation is already available in Spark ML.”
      ↩︎ Removing duplicate and near-duplicate documents
      “You can eliminate some duplicates using metadata alone.”
      ↩︎ Exam trap 3
      “Filtering: Eliminate irrelevant or unwanted documents from the collection.”
      ↩︎ Exam trap 4
      “removing duplicates, and filtering out unwanted information, the next step is to break it down into smaller, manageable units called chunks.”
      ↩︎ Checkpoint
      “applying a personally identifiable information (PII) detection algorithm to the documents to filter documents.”
      ↩︎ Checkpoint
      “Fit a MinHash model and hash the vectors using MinHash for Jaccard distance.”
      ↩︎ Checkpoint
    3. 3.
      “Remove boilerplate (headers, footers, page numbers).”
      ↩︎ Cleaning boilerplate and noise during parsing
      “Poor parsing (missing tables, broken formatting) directly impacts retrieval quality.”
      ↩︎ Cleaning boilerplate and noise during parsing
      “Semantic metadata: Add document summaries and section headers to chunks.”
      ↩︎ Separating layout markers from body text with ai_parse_document
      “Preserve document structure (headings, lists, tables).”
      ↩︎ Checkpoint
    4. 4.
      “identifies and extracts layout information from a document, like page numbers, headers, tables, and footers, and returns them as structured elements.”
      ↩︎ Separating layout markers from body text with ai_parse_document
      “a layout marker like a page header or footer.”
      ↩︎ Separating layout markers from body text with ai_parse_document
      “section_header: A heading or subheading that denotes the start of a section.”
      ↩︎ Exam trap 2
      “page_header: A header that appears at the top of a page. page_footer: A footer that appears at the bottom of a page.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.