What you will be able to do
- Pick a parsing package based on the source format: unstructured or PyPDF2 for PDFs and Word documents, BeautifulSoup or lxml for HTML, an OCR library for scanned images, and ai_parse_document for layout-aware parsing on Databricks
- Load raw files as bytes with the binaryFile reader so any of these parsers can use them
- Build a PyPDF2 UDF to use when AI Functions are not available
- Use ai_parse_document's options, element types, VARIANT output and limits to decide whether it suits a given workload
Key concept
Format-driven parser selection — There is no single best extraction package. You choose the parser to match the format of the source files and the structure you need to keep, such as plain text, HTML tags, OCR for images, or layout elements like tables and headers.
1.Start from the format: which package handles which source
In a RAG data pipeline, parsing is the first preprocessing step after ingestion. It pulls the content you need out of raw files and turns it into something you can chunk and embed. The Databricks RAG cookbook doesn't name one universal tool. It ties each kind of source to a family of libraries. On the exam, questions in this area usually give you a source format and ask which package fits it, so learn the mapping first.
The cookbook covers three groups of formats. For text documents such as PDFs and Word files, it names the general-purpose libraries unstructured and PyPDF2. These cover several file formats and let you customize how parsing works. For HTML, it names BeautifulSoup and lxml, which can navigate the HTML structure, select specific elements, and extract the text or attributes you want. For images and scanned documents, it says you usually need Optical Character Recognition (OCR). Examples are the open-source Tesseract and managed services such as Amazon Textract, Azure AI Vision OCR and Google Cloud Vision API.
Databricks also has a managed option, the ai_parse_document AI Function. It accepts PDF, JPG/JPEG, PNG, TIFF/TIF, DOC/DOCX and PPT/PPTX files and returns layout-aware structured elements instead of a flat string. HTML is not on its supported-format list.
| Source format | Package(s) named | What you get back |
|---|---|---|
| PDFs, Word docs | unstructured, PyPDF2 | Extracted text, with options to customize parsing |
| HTML web pages | BeautifulSoup, lxml | Selected elements, text or attributes from the HTML structure |
| Images and scanned documents | OCR: Tesseract, Amazon Textract, Azure AI Vision OCR, Google Cloud Vision API | Text recognized from the image |
| Format family | Supported extensions |
|---|---|
| Documents | PDF, DOC/DOCX |
| Presentations | PPT/PPTX |
| Images | JPG/JPEG, PNG, TIFF/TIF |
Checkpoint 1 of 9· Match them up
Match each package to the source it is named for in the Databricks RAG guidance
Tap a term, then the definition that fits it.
The cookbook assigns general-purpose document libraries to PDFs and Word files, HTML parsers to web pages, and OCR libraries to images and scans.
“Off-the-shelf libraries like unstructured and PyPDF2 can handle various file formats and provide options for customizing the parsing process.”Source: docs.databricks.com
Checkpoint 2 of 9· Exam question
A team is building a RAG corpus from PDF product manuals that contain large multi-page tables needing to stay structured, plus several pages that are scanned images requiring text recognition. Which approach best extracts this content?
Correct answer: B — The `ai_parse_document` SQL function calls a Databricks Foundation Model to extract layout, HTML-formatted tables, and OCR text from scanned pages in a single pass.
- A. BeautifulSoup and lxml are built to navigate HTML's tag-based DOM, not the binary page-stream layout of a PDF, so it cannot select tables or paragraphs out of a PDF file.
- B. This is correct: ai_parse_document runs inside the Lakehouse and uses a Databricks Foundation Model to jointly extract layout, render tables as HTML, and OCR scanned pages in one function call.
- C. Converting a PDF to .docx does not automatically infer table boundaries or perform OCR on scanned images; python-docx only reads structure that already exists in a Word file's XML.
- D. pypdf extracts the embedded text layer of a born-digital PDF but has no OCR capability and does not reconstruct table structure, so scanned pages and complex tables come out as unstructured or missing text.
Sources1
2.Every parser starts from bytes: the binaryFile reader
Whichever package you choose, on Databricks the raw documents reach it the same way. Spark's binaryFile data source reads each file as a single record. That record holds the file's raw content plus its metadata. The resulting DataFrame has four columns: path (string), modificationTime (timestamp), length (long) and content (binary). Your parser works on the content column.
For ai_parse_document, this input format is required: the input must be a binary column in a DataFrame or Delta table. If the files live in a Unity Catalog volume, the binaryFile reader is how you produce that column. To read only some formats, pathGlobFilter limits the files to names that match a pattern such as *.jpg, and recursiveFileLookup searches nested directories.
Checkpoint 3 of 9· Fill the gap
Which data source format loads every file in a volume as one row with a binary content column?
df = spark.read.format(" ? ").load("/Volumes/<catalog>/<schema>/<volume>/")binaryFile turns each file into one record with path, modificationTime, length and a BinaryType content column, which is the input both PyPDF2 UDFs and ai_parse_document expect.
Source: docs.databricks.compathGlobFilter, for example .option("pathGlobFilter", "*.pdf"). It loads only files whose names match the pattern. Partition discovery still applies.
3.The open-source path: a PyPDF2 UDF
Databricks AI Functions need a workspace in a supported region. For workspaces without access, the Databricks unstructured-data tutorial gives an alternative that uses standard Python libraries. For PDFs, you install a pinned PyPDF2 (%pip install PyPDF2==3.0.1) and wrap its PdfReader in a Spark UDF. The UDF gets the bytes from the content column, wraps them in io.BytesIO, and joins the text extracted from every page. The tutorial's image example works the same way: it opens the bytes with Pillow to read width, height and format.
from PyPDF2 import PdfReader
import io
@udf(returnType=StringType())
def extract_pdf_text(content):
if content is None:
return None
try:
reader = PdfReader(io.BytesIO(content))
return "\n".join(page.extract_text() or "" for page in reader.pages)
except Exception as e:
return f"Error: {str(e)}"
df = spark.read.format("binaryFile") \
.option("pathGlobFilter", "*.pdf") \
.load("/Volumes/unstructured_data_lab/raw/files_volume/")
result_df = df.withColumn("text_content", extract_pdf_text("content"))
display(result_df.select("path", "text_content"))The output is one plain string per PDF. That is the main difference from the managed path below. The page text is concatenated, and nothing labels which part was a table, a heading or a footer. When the source documents are simple text, that may be enough. When they are layout-heavy, you have to write any structure recovery or cleanup yourself.
Checkpoint 4 of 9· Check yourself
Your workspace is in a region where AI Functions are unavailable, and you must extract text from a volume of PDFs. Which approach does Databricks documentation give for this case?
The REST API is the same AI Function and is subject to the same regional availability. The tutorial's fallback for this case is a Python library, PyPDF2, applied in a UDF.
“If AI functions aren't available in your region, use Python libraries:”Source: docs.databricks.com
Checkpoint 5 of 9· Exam question
A developer is scraping crawled HTML pages from a documentation website and needs to strip navigation bars, headers, and footers while selecting only the article body using specific CSS classes and tag hierarchy. Which package fits this need?
Correct answer: A — Parsing the pages with `BeautifulSoup` and the `lxml` parser navigates the DOM tree and uses `.select()` or `.find()` to pull elements by tag, id, and CSS class.
- A. This is correct: BeautifulSoup with the lxml parser is purpose-built to walk an HTML DOM and target elements by tag, id, or CSS class, which is exactly the fine-grained selection this scraping task needs.
- B. pypdf reads PDF page-stream content, not HTML markup, so it has no concept of tags, classes, or a DOM tree to navigate.
- C. python-docx reads the XML parts of a Word .docx file and has no support for parsing HTML markup or CSS selectors.
- D. The unstructured library's generic partitioner returns normalized elements like titles and narrative text, but it does not expose the underlying CSS class or id attributes needed for targeted element selection.
Sources4
4.The managed path: ai_parse_document and its structured output
ai_parse_document runs on a model that Databricks manages and serves through Foundation Model APIs. It first identifies layout, such as page numbers, headers, tables and footers, and then extracts the content inside each region. The result is a sequence of elements. Each element has an id, a type, the extracted content, a confidence score, a bbox location, and an optional AI-generated description. The element types are text, table, figure, title, caption, section_header, page_header, page_footer, page_number and footnote. Because page headers and footers come back labeled as their own types, you can filter them out later instead of finding them by string matching.
You can call the function from SQL, from PySpark as dbf.ai_parse_document(col=<col>, options=<options>), or in Scala. It requires Databricks Runtime 17.3 or above. On serverless compute, the environment version must be 3 or above. The only required argument is content. The optional map controls the rest.
| Option | Effect | When to set it |
|---|---|---|
| version | Pins the output schema version; "2.0" is supported | To lock the schema your downstream code reads |
| imageOutputPath | Saves rendered page images to a Unity Catalog volume | Reference images or multi-modal RAG |
| descriptionElementTypes | Controls AI-generated descriptions; '' turns them off, 'figure' or '*' describes figures | Set '' to cut compute and cost on figure-heavy documents |
| pageRange | Parses only listed pages or ranges, for example '1,3,5-10' | Documents over 500 pages, which otherwise fail immediately |
df = spark.read.format("binaryFile") \
.load("/Volumes/path/to/your/directory") \
.withColumn(
"parsed",
expr("ai_parse_document(content)"))
display(df)The output type has a practical consequence. The result is a VARIANT, and PySpark can't collect a VARIANT directly. To work with parsed documents as Python objects, convert the result to a JSON string with to_json() in SQL, collect the rows, and call json.loads() on each string. In SQL, you can also split the top-level fields into separate columns with paths such as parsed:document:elements and parsed:error_status.
Checkpoint 6 of 9· Put it in order
Put these steps in order to bring ai_parse_document results into Python dictionaries
- 1.Convert the VARIANT result to a JSON string with to_json() in SQL
- 2.Collect the rows and parse each string with json.loads() in Python
- 3.Apply ai_parse_document to the content column
- 4.Read the source files with format binaryFile
The function needs binary input and returns a VARIANT. Because PySpark can't collect a VARIANT, you serialize it with to_json() before calling json.loads().
“use to_json() in SQL to convert the VARIANT to a JSON string, then parse it with json.loads() in Python”Source: docs.databricks.com
Before choosing the managed path, check its limits. Documents can have at most 500 pages and files can be at most 100 MB. REST API requests are limited to 100 pages per document. The function is available only in some regions. You can't customize the model or bring your own. Results may be weaker for images containing non-Latin text such as Japanese or Korean, and documents with digital signatures may not be processed accurately. To check results before scaling up, the Document Parsing UI in Agent Bricks shows each source document next to its parsed output.
Checkpoint 7 of 9· Check yourself
Which workload is the weakest fit for ai_parse_document as documented?
The documentation states that the model may not perform well on images with non-Latin text. The 900-page PDF works because pageRange keeps the selection within the 500-page limit.
“The underlying model may not perform optimally when handling images using text of non-Latin alphabets, such as Japanese or Korean.”Source: docs.databricks.com
Checkpoint 8 of 9· Exam question
An internal knowledge base consists exclusively of Microsoft Word `.docx` files with heading styles and paragraph formatting that must be preserved for structure-aware chunking. Which package should the engineer use to extract this content?
Correct answer: D — Opening each file with `python-docx` reads paragraphs and their applied heading styles directly from the document's XML parts, preserving the structural hierarchy.
- A. pypdf is built to read PDF page streams, not the Open XML package format that .docx files use, so it cannot parse Word paragraph or heading style metadata.
- B. A .docx file is not HTML, so BeautifulSoup has no tag structure to navigate; treating it as an HTML document would not expose Word's heading styles.
- C. A .docx file stores text and styles as XML, not as page images, so there is nothing for an OCR engine like Tesseract to recognize, and heading hierarchy would be lost entirely.
- D. This is correct: python-docx reads a .docx file's underlying XML parts directly, giving access to each paragraph's text and its applied heading style without any image rendering or OCR step.
5.Whichever package you pick: parsing practices that protect RAG quality
Choosing a package is half the job. The cookbook's parsing best practices apply whatever you choose. First, clean the extracted text to remove noise such as headers, footers and special characters, so the RAG chain handles less malformed input. Second, add error handling and logging. The PyPDF2 UDF above returns an Error: string instead of failing the job, and those errors often point to problems in the source data. Third, be prepared to customize parsing logic for your document structure. The cookbook says this upfront effort often prevents downstream quality problems. Fourth, review a sample of the parsed output by hand on a regular basis. For ai_parse_document, the Document Parsing UI supports this check, because it shows each document region next to the content extracted from it.
Checkpoint 9 of 9· Check yourself
Your parsing step logs errors for a growing share of files. According to the cookbook, what do such parsing errors often indicate?
The cookbook recommends error handling and logging during parsing because the errors it surfaces often trace back to problems in the source data, not to later pipeline stages.
“Doing so often points to upstream issues with the quality of the source data.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.ai_parse_document is the right answer for any source format, including HTML web pages.Why is that wrong?
ai_parse_document supports PDF, image, DOC/DOCX and PPT/PPTX files. HTML is not on that list, and the cookbook points HTML to BeautifulSoup or lxml.
Covered in Start from the format: which package handles which source
2.If AI Functions aren't available, calling ai_parse_document through the REST API gets around the restriction.Why is that wrong?
The REST API calls the same AI Function, which is available only in some regions. The documented fallback is a Python library such as PyPDF2 applied in a UDF.
Covered in The open-source path: a PyPDF2 UDF
3.You can call .collect() on the ai_parse_document column and get Python dictionaries straight away.Why is that wrong?
The function returns a VARIANT, which PySpark can't collect directly. Convert it with to_json() first, then call json.loads().
Covered in The managed path: ai_parse_document and its structured output
4.A document over 500 pages is parsed up to page 500 and the rest is silently skipped.Why is that wrong?
Without pageRange, the call fails before parsing any page. Use pageRange to select a subset that stays within the limit.
Covered in The managed path: ai_parse_document and its structured output
Practise it for real
Extract text from the PDFs in a Unity Catalog volume with PyPDF2, without using AI Functions
1.In a notebook, run %pip install PyPDF2==3.0.1
Why: The tutorial pins this version for the PdfReader-based fallback
You should see: PyPDF2 installs into the notebook environment
2.Define the extract_pdf_text UDF from the lesson, which wraps the content bytes in io.BytesIO and passes them to PdfReader
Why: The UDF turns each binary content value into one text string and returns an Error: string instead of failing
You should see: A UDF that returns StringType
3.Read the volume with spark.read.format("binaryFile").option("pathGlobFilter", "*.pdf")
Why: binaryFile gives the content column the UDF needs, and the glob keeps the read to PDFs
You should see: A DataFrame with path, modificationTime, length and content columns
4.Add the text_content column with withColumn and display path and text_content
Why: This lets you review a sample of the parsed output by hand, as the cookbook recommends
You should see: One row per PDF, with its page text joined by newlines or an Error: message
Stuck? Get a nudge
If text_content is empty for some files, they may be scanned images. The cookbook routes those to OCR.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“Off-the-shelf libraries like unstructured and PyPDF2 can handle various file formats and provide options for customizing the parsing process.”
↩︎ Start from the format: which package handles which source“such as Tesseract or SaaS versions like Amazon Textract, Azure AI Vision OCR, and Google Cloud Vision API”
↩︎ Start from the format: which package handles which source“Optical Character Recognition (OCR) techniques are typically required to extract text from images.”
↩︎ Start from the format: which package handles which source“Data cleaning: Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters.”
↩︎ Whichever package you pick: parsing practices that protect RAG quality“Evaluating parsing quality: Regularly assess the quality of the parsed data by manually reviewing a sample of the output.”
↩︎ Whichever package you pick: parsing practices that protect RAG quality“The specific parsing techniques and tools you use depend on the type of data you are working with.”
↩︎ Key concept“HTML parsing libraries like BeautifulSoup and lxml can be used to extract relevant content from web pages.”
↩︎ Exam trap 1“HTML parsing libraries like BeautifulSoup and lxml can be used to extract relevant content from web pages.”
↩︎ Prediction“Doing so often points to upstream issues with the quality of the source data.”
↩︎ Checkpoint - 2.
“reads binary files and converts each file into a single record containing the file's raw content and metadata”
↩︎ Every parser starts from bytes: the binaryFile reader - 3.
“Your input data files must be stored as blob data in bytes, meaning a binary type column in a DataFrame or Delta table.”
↩︎ Every parser starts from bytes: the binaryFile reader“identifies and extracts layout information from a document, like page numbers, headers, tables, and footers, and returns them as structured elements.”
↩︎ The managed path: ai_parse_document and its structured output“Documents are limited to a maximum of 500 pages, exceeding this limit results in errors.”
↩︎ The managed path: ai_parse_document and its structured output“It renders the source document alongside the parsed output, letting you inspect what content was extracted from each region of your documents.”
↩︎ Whichever package you pick: parsing practices that protect RAG quality“ai_parse_document returns a VARIANT type, which cannot be directly collected by PySpark”
↩︎ Exam trap 3“If pageRange is omitted and the document exceeds 500 pages, the function fails immediately without parsing any pages.”
↩︎ Exam trap 4“For version 2.0, tables are represented in HTML. The output is of VARIANT type.”
↩︎ Prediction“use to_json() in SQL to convert the VARIANT to a JSON string, then parse it with json.loads() in Python”
↩︎ Checkpoint“The underlying model may not perform optimally when handling images using text of non-Latin alphabets, such as Japanese or Korean.”
↩︎ Checkpoint - 4.
“If AI functions aren't available in your region, use Python libraries:”
↩︎ The open-source path: a PyPDF2 UDF“If you don't have access to AI functions, use standard Python libraries instead.”
↩︎ Exam trap 2 - 5.
“Parses a column containing binary data (blob) and returns a VariantType.”
↩︎ The managed path: ai_parse_document and its structured output - 6.
“ai_parse_document REST API requests are limited to 100 pages per document.”
↩︎ The managed path: ai_parse_document and its structured output