CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 9/56

    Choosing a Python Package to Parse PDF, HTML, DOCX and Images

    Choose the appropriate Python package to extract document content from provided source data and format.

    16 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Pick a parsing package based on the source format: unstructured or PyPDF2 for PDFs and Word documents, BeautifulSoup or lxml for HTML, an OCR library for scanned images, and ai_parse_document for layout-aware parsing on Databricks
    • Load raw files as bytes with the binaryFile reader so any of these parsers can use them
    • Build a PyPDF2 UDF to use when AI Functions are not available
    • Use ai_parse_document's options, element types, VARIANT output and limits to decide whether it suits a given workload

    Key concept

    Format-driven parser selection — There is no single best extraction package. You choose the parser to match the format of the source files and the structure you need to keep, such as plain text, HTML tags, OCR for images, or layout elements like tables and headers.

    1.Start from the format: which package handles which source

    In a RAG data pipeline, parsing is the first preprocessing step after ingestion. It pulls the content you need out of raw files and turns it into something you can chunk and embed. The Databricks RAG cookbook doesn't name one universal tool. It ties each kind of source to a family of libraries. On the exam, questions in this area usually give you a source format and ask which package fits it, so learn the mapping first.

    The cookbook covers three groups of formats. For text documents such as PDFs and Word files, it names the general-purpose libraries unstructured and PyPDF2. These cover several file formats and let you customize how parsing works. For HTML, it names BeautifulSoup and lxml, which can navigate the HTML structure, select specific elements, and extract the text or attributes you want. For images and scanned documents, it says you usually need Optical Character Recognition (OCR). Examples are the open-source Tesseract and managed services such as Amazon Textract, Azure AI Vision OCR and Google Cloud Vision API.

    Databricks also has a managed option, the ai_parse_document AI Function. It accepts PDF, JPG/JPEG, PNG, TIFF/TIF, DOC/DOCX and PPT/PPTX files and returns layout-aware structured elements instead of a flat string. HTML is not on its supported-format list.

    Source format and the packages the Databricks RAG guidance names for it
    Source formatPackage(s) namedWhat you get back
    PDFs, Word docsunstructured, PyPDF2Extracted text, with options to customize parsing
    HTML web pagesBeautifulSoup, lxmlSelected elements, text or attributes from the HTML structure
    Images and scanned documentsOCR: Tesseract, Amazon Textract, Azure AI Vision OCR, Google Cloud Vision APIText recognized from the image
    Formats accepted by the Databricks-managed ai_parse_document function
    Format familySupported extensions
    DocumentsPDF, DOC/DOCX
    PresentationsPPT/PPTX
    ImagesJPG/JPEG, PNG, TIFF/TIF

    Checkpoint 1 of 9· Match them up

    Match each package to the source it is named for in the Databricks RAG guidance

    Tap a term, then the definition that fits it.

    Checkpoint 2 of 9· Exam question

    A team is building a RAG corpus from PDF product manuals that contain large multi-page tables needing to stay structured, plus several pages that are scanned images requiring text recognition. Which approach best extracts this content?

    Sources1

    2.Every parser starts from bytes: the binaryFile reader

    Whichever package you choose, on Databricks the raw documents reach it the same way. Spark's binaryFile data source reads each file as a single record. That record holds the file's raw content plus its metadata. The resulting DataFrame has four columns: path (string), modificationTime (timestamp), length (long) and content (binary). Your parser works on the content column.

    For ai_parse_document, this input format is required: the input must be a binary column in a DataFrame or Delta table. If the files live in a Unity Catalog volume, the binaryFile reader is how you produce that column. To read only some formats, pathGlobFilter limits the files to names that match a pattern such as *.jpg, and recursiveFileLookup searches nested directories.

    Checkpoint 3 of 9· Fill the gap

    Which data source format loads every file in a volume as one row with a binary content column?

    df = spark.read.format(" ? ").load("/Volumes/<catalog>/<schema>/<volume>/")

    Sources23

    3.The open-source path: a PyPDF2 UDF

    Databricks AI Functions need a workspace in a supported region. For workspaces without access, the Databricks unstructured-data tutorial gives an alternative that uses standard Python libraries. For PDFs, you install a pinned PyPDF2 (%pip install PyPDF2==3.0.1) and wrap its PdfReader in a Spark UDF. The UDF gets the bytes from the content column, wraps them in io.BytesIO, and joins the text extracted from every page. The tutorial's image example works the same way: it opens the bytes with Pillow to read width, height and format.

    Extracting PDF text with PyPDF2 inside a Spark UDF when AI Functions aren't availablepython
    from PyPDF2 import PdfReader
    import io
    
    @udf(returnType=StringType())
    def extract_pdf_text(content):
        if content is None:
            return None
        try:
            reader = PdfReader(io.BytesIO(content))
            return "\n".join(page.extract_text() or "" for page in reader.pages)
        except Exception as e:
            return f"Error: {str(e)}"
    
    df = spark.read.format("binaryFile") \
        .option("pathGlobFilter", "*.pdf") \
        .load("/Volumes/unstructured_data_lab/raw/files_volume/")
    
    result_df = df.withColumn("text_content", extract_pdf_text("content"))
    display(result_df.select("path", "text_content"))

    The output is one plain string per PDF. That is the main difference from the managed path below. The page text is concatenated, and nothing labels which part was a table, a heading or a footer. When the source documents are simple text, that may be enough. When they are layout-heavy, you have to write any structure recovery or cleanup yourself.

    Checkpoint 4 of 9· Check yourself

    Your workspace is in a region where AI Functions are unavailable, and you must extract text from a volume of PDFs. Which approach does Databricks documentation give for this case?

    Checkpoint 5 of 9· Exam question

    A developer is scraping crawled HTML pages from a documentation website and needs to strip navigation bars, headers, and footers while selecting only the article body using specific CSS classes and tag hierarchy. Which package fits this need?

    Sources4

    4.The managed path: ai_parse_document and its structured output

    ai_parse_document runs on a model that Databricks manages and serves through Foundation Model APIs. It first identifies layout, such as page numbers, headers, tables and footers, and then extracts the content inside each region. The result is a sequence of elements. Each element has an id, a type, the extracted content, a confidence score, a bbox location, and an optional AI-generated description. The element types are text, table, figure, title, caption, section_header, page_header, page_footer, page_number and footnote. Because page headers and footers come back labeled as their own types, you can filter them out later instead of finding them by string matching.

    You can call the function from SQL, from PySpark as dbf.ai_parse_document(col=<col>, options=<options>), or in Scala. It requires Databricks Runtime 17.3 or above. On serverless compute, the environment version must be 3 or above. The only required argument is content. The optional map controls the rest.

    Optional settings in the ai_parse_document options map
    OptionEffectWhen to set it
    versionPins the output schema version; "2.0" is supportedTo lock the schema your downstream code reads
    imageOutputPathSaves rendered page images to a Unity Catalog volumeReference images or multi-modal RAG
    descriptionElementTypesControls AI-generated descriptions; '' turns them off, 'figure' or '*' describes figuresSet '' to cut compute and cost on figure-heavy documents
    pageRangeParses only listed pages or ranges, for example '1,3,5-10'Documents over 500 pages, which otherwise fail immediately
    Parsing every file in a volume with ai_parse_document from PySparkpython
    df = spark.read.format("binaryFile") \
      .load("/Volumes/path/to/your/directory") \
      .withColumn(
        "parsed",
        expr("ai_parse_document(content)"))
    display(df)

    The output type has a practical consequence. The result is a VARIANT, and PySpark can't collect a VARIANT directly. To work with parsed documents as Python objects, convert the result to a JSON string with to_json() in SQL, collect the rows, and call json.loads() on each string. In SQL, you can also split the top-level fields into separate columns with paths such as parsed:document:elements and parsed:error_status.

    Checkpoint 6 of 9· Put it in order

    Put these steps in order to bring ai_parse_document results into Python dictionaries

    1. 1.Convert the VARIANT result to a JSON string with to_json() in SQL
    2. 2.Collect the rows and parse each string with json.loads() in Python
    3. 3.Apply ai_parse_document to the content column
    4. 4.Read the source files with format binaryFile

    Before choosing the managed path, check its limits. Documents can have at most 500 pages and files can be at most 100 MB. REST API requests are limited to 100 pages per document. The function is available only in some regions. You can't customize the model or bring your own. Results may be weaker for images containing non-Latin text such as Japanese or Korean, and documents with digital signatures may not be processed accurately. To check results before scaling up, the Document Parsing UI in Agent Bricks shows each source document next to its parsed output.

    Checkpoint 7 of 9· Check yourself

    Which workload is the weakest fit for ai_parse_document as documented?

    Checkpoint 8 of 9· Exam question

    An internal knowledge base consists exclusively of Microsoft Word `.docx` files with heading styles and paragraph formatting that must be preserved for structure-aware chunking. Which package should the engineer use to extract this content?

    Sources356

    5.Whichever package you pick: parsing practices that protect RAG quality

    Choosing a package is half the job. The cookbook's parsing best practices apply whatever you choose. First, clean the extracted text to remove noise such as headers, footers and special characters, so the RAG chain handles less malformed input. Second, add error handling and logging. The PyPDF2 UDF above returns an Error: string instead of failing the job, and those errors often point to problems in the source data. Third, be prepared to customize parsing logic for your document structure. The cookbook says this upfront effort often prevents downstream quality problems. Fourth, review a sample of the parsed output by hand on a regular basis. For ai_parse_document, the Document Parsing UI supports this check, because it shows each document region next to the content extracted from it.

    Checkpoint 9 of 9· Check yourself

    Your parsing step logs errors for a growing share of files. According to the cookbook, what do such parsing errors often indicate?

    Sources13

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.ai_parse_document is the right answer for any source format, including HTML web pages.Why is that wrong?

      ai_parse_document supports PDF, image, DOC/DOCX and PPT/PPTX files. HTML is not on that list, and the cookbook points HTML to BeautifulSoup or lxml.

      Covered in Start from the format: which package handles which source

    2. 2.If AI Functions aren't available, calling ai_parse_document through the REST API gets around the restriction.Why is that wrong?

      The REST API calls the same AI Function, which is available only in some regions. The documented fallback is a Python library such as PyPDF2 applied in a UDF.

      Covered in The open-source path: a PyPDF2 UDF

    3. 3.You can call .collect() on the ai_parse_document column and get Python dictionaries straight away.Why is that wrong?

      The function returns a VARIANT, which PySpark can't collect directly. Convert it with to_json() first, then call json.loads().

      Covered in The managed path: ai_parse_document and its structured output

    4. 4.A document over 500 pages is parsed up to page 500 and the rest is silently skipped.Why is that wrong?

      Without pageRange, the call fails before parsing any page. Use pageRange to select a subset that stays within the limit.

      Covered in The managed path: ai_parse_document and its structured output

    Practise it for real

    Extract text from the PDFs in a Unity Catalog volume with PyPDF2, without using AI Functions

    1. 1.In a notebook, run %pip install PyPDF2==3.0.1

      Why: The tutorial pins this version for the PdfReader-based fallback

      You should see: PyPDF2 installs into the notebook environment

    2. 2.Define the extract_pdf_text UDF from the lesson, which wraps the content bytes in io.BytesIO and passes them to PdfReader

      Why: The UDF turns each binary content value into one text string and returns an Error: string instead of failing

      You should see: A UDF that returns StringType

    3. 3.Read the volume with spark.read.format("binaryFile").option("pathGlobFilter", "*.pdf")

      Why: binaryFile gives the content column the UDF needs, and the glob keeps the read to PDFs

      You should see: A DataFrame with path, modificationTime, length and content columns

    4. 4.Add the text_content column with withColumn and display path and text_content

      Why: This lets you review a sample of the parsed output by hand, as the cookbook recommends

      You should see: One row per PDF, with its page text joined by newlines or an Error: message

    Stuck? Get a nudge

    If text_content is empty for some files, they may be scanned images. The cookbook routes those to OCR.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Off-the-shelf libraries like unstructured and PyPDF2 can handle various file formats and provide options for customizing the parsing process.”
      ↩︎ Start from the format: which package handles which source
      “such as Tesseract or SaaS versions like Amazon Textract, Azure AI Vision OCR, and Google Cloud Vision API”
      ↩︎ Start from the format: which package handles which source
      “Optical Character Recognition (OCR) techniques are typically required to extract text from images.”
      ↩︎ Start from the format: which package handles which source
      “Data cleaning: Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters.”
      ↩︎ Whichever package you pick: parsing practices that protect RAG quality
      “Evaluating parsing quality: Regularly assess the quality of the parsed data by manually reviewing a sample of the output.”
      ↩︎ Whichever package you pick: parsing practices that protect RAG quality
      “The specific parsing techniques and tools you use depend on the type of data you are working with.”
      ↩︎ Key concept
      “HTML parsing libraries like BeautifulSoup and lxml can be used to extract relevant content from web pages.”
      ↩︎ Exam trap 1
      “HTML parsing libraries like BeautifulSoup and lxml can be used to extract relevant content from web pages.”
      ↩︎ Prediction
      “Doing so often points to upstream issues with the quality of the source data.”
      ↩︎ Checkpoint
    2. 2.
      “reads binary files and converts each file into a single record containing the file's raw content and metadata”
      ↩︎ Every parser starts from bytes: the binaryFile reader
    3. 3.
      “Your input data files must be stored as blob data in bytes, meaning a binary type column in a DataFrame or Delta table.”
      ↩︎ Every parser starts from bytes: the binaryFile reader
      “identifies and extracts layout information from a document, like page numbers, headers, tables, and footers, and returns them as structured elements.”
      ↩︎ The managed path: ai_parse_document and its structured output
      “Documents are limited to a maximum of 500 pages, exceeding this limit results in errors.”
      ↩︎ The managed path: ai_parse_document and its structured output
      “It renders the source document alongside the parsed output, letting you inspect what content was extracted from each region of your documents.”
      ↩︎ Whichever package you pick: parsing practices that protect RAG quality
      “ai_parse_document returns a VARIANT type, which cannot be directly collected by PySpark”
      ↩︎ Exam trap 3
      “If pageRange is omitted and the document exceeds 500 pages, the function fails immediately without parsing any pages.”
      ↩︎ Exam trap 4
      “For version 2.0, tables are represented in HTML. The output is of VARIANT type.”
      ↩︎ Prediction
      “use to_json() in SQL to convert the VARIANT to a JSON string, then parse it with json.loads() in Python”
      ↩︎ Checkpoint
      “The underlying model may not perform optimally when handling images using text of non-Latin alphabets, such as Japanese or Korean.”
      ↩︎ Checkpoint
    4. 4.
      “If AI functions aren't available in your region, use Python libraries:”
      ↩︎ The open-source path: a PyPDF2 UDF
      “If you don't have access to AI functions, use standard Python libraries instead.”
      ↩︎ Exam trap 2
    5. 6.
      “ai_parse_document REST API requests are limited to 100 pages per document.”
      ↩︎ The managed path: ai_parse_document and its structured output

    Spotted a mistake, or was something unclear? Tell us.