What you will be able to do
- Write an AI_PARSE_DOCUMENT call against a staged file using TO_FILE and an options object
- Choose between OCR and LAYOUT mode for a given document and know what each returns
- Predict how page_split changes the shape of the JSON output and which file formats it accepts
- Restrict parsing to specific pages with page_filter and state the per-document page ceiling
- Describe what AI_PARSE_DOCUMENT returns when a document cannot be parsed
Key concept
Parsing mode decides the output shape — AI_PARSE_DOCUMENT turns a staged document into a JSON string. The options object controls what comes back: OCR mode returns plain text, LAYOUT mode returns Markdown that keeps the document's structure, and page options decide whether you get one content field or a pages array.
1.The shape of an AI_PARSE_DOCUMENT call
AI_PARSE_DOCUMENT is the Cortex AI function that turns a document into text Snowflake can work with. It reads files on internal or external stages and keeps the reading order and structural elements such as tables and headers. Its output can then feed other Cortex AI functions for extraction, classification or enrichment. The function takes one required argument and two optional ones:
AI_PARSE_DOCUMENT( <file_object> [, <options> ] [, <return_error_details> ] )The first argument is a FILE object pointing to a document on a Snowflake stage, normally built with TO_FILE('@stage', 'path'). The second is an OBJECT of options. Every key in it is optional, and the keys you will meet most are 'mode', 'page_split', 'page_filter' and 'extract_images'. The function returns a JSON-formatted string, so the documentation suggests passing it through PARSE_JSON before using the result in SQL.
The older function, SNOWFLAKE.CORTEX.PARSE_DOCUMENT, was called differently: it took a stage name and a relative path as two separate string arguments. Its reference page is kept only for backward compatibility and points new work to AI_PARSE_DOCUMENT.
Checkpoint 1 of 8· Check yourself
You are writing new code to parse a PDF staged at @docs.doc_stage. Which call matches the current function?
AI_PARSE_DOCUMENT takes a FILE object, usually built with TO_FILE, followed by an optional options OBJECT. The stage-plus-path string form belongs to the legacy PARSE_DOCUMENT function.
“A FILE object that specifies the document to parse, stored in a Snowflake stage.”Source: docs.snowflake.com
2.OCR mode vs LAYOUT mode
The 'mode' key picks one of two extraction types. OCR extracts text only and is the default. LAYOUT extracts layout as well as text, including structural content such as tables. The two modes also return different formats: content is plain text in OCR mode and Markdown in LAYOUT mode. Headings come back as # lines, and tables come back as Markdown tables.
The defaults and the recommendations point in different directions. The user guide calls LAYOUT the preferred choice for most use cases, especially complex documents, and it is required for image extraction. It recommends OCR for fast, high-quality text extraction from scanned or text-heavy documents such as contracts, insurance claims and manuals. So if you want tables or headings to survive, you have to ask for LAYOUT explicitly.
| Aspect | OCR | LAYOUT |
|---|---|---|
| What is extracted | Text only | Text plus layout, including tables and headers |
| Format of content | Plain text | Markdown-formatted text |
| Default? | Yes | No: set 'mode': 'LAYOUT' |
| Recommended for | Fast text extraction from scanned or text-heavy documents | Most use cases, especially complex documents |
| 'extract_images' | Not available | Available, at no additional cost |
| Document Processing Playground tab | Text | Markdown |
Image extraction is only available in LAYOUT. Setting 'extract_images' to TRUE adds an "images" array to the output. Each entry has an id, bounding-box coordinates (top_left_x, top_left_y, bottom_right_x, bottom_right_y) and the image itself as image_base64. The extracted images can then go to AI_EXTRACT or AI_COMPLETE for tagging and analysis.
Checkpoint 2 of 8· Match them up
Match each option or mode to what it produces
Tap a term, then the definition that fits it.
OCR returns plain text and LAYOUT returns Markdown. Image extraction depends on LAYOUT, and page_split changes the output into a per-page array.
“Plain text (in OCR mode) or Markdown-formatted text (in LAYOUT mode).”Source: docs.snowflake.com
Checkpoint 3 of 8· Exam question
A claims processing team receives scanned insurance claim PDFs that are dense with body text and have almost no tables. They need AI_PARSE_DOCUMENT to extract the text quickly and cost-effectively for downstream analysis. Which configuration best fits this requirement?
Correct answer: A — Call AI_PARSE_DOCUMENT with mode set to OCR, since it is built for fast, high-quality text extraction from scanned, text-heavy documents like claims
- A. OCR mode is recommended for fast, high-quality text extraction from scanned or text-heavy documents such as insurance claims, which matches this workload exactly.
- B. LAYOUT mode is optimized for preserving structural content like tables and headings, which is unnecessary overhead when the claim documents are mostly dense body text.
- C. Extracting embedded images is not the stated need, and the function's default mode is OCR rather than LAYOUT, so this description misstates the baseline behavior.
- D. The legacy PARSE_DOCUMENT function is being superseded by AI_PARSE_DOCUMENT and is not documented as offering better throughput; it is treated as the older form of the current function.
3.page_split: one result per page
By default, AI_PARSE_DOCUMENT returns the whole document as a single "content" field. When 'page_split' is TRUE, the document is split into pages and each page is processed separately. The output then becomes a "pages" array, and each element has a "content" field and an "index" field. The index starts at 0, and any page numbering printed in the document is ignored. A one-page document still comes back as a "pages" array with one element, so downstream code can rely on the same shape every time.
This is the documented way to handle long documents. The reference tip says to set the option to TRUE for documents that exceed the function's token limit. The user guide's examples combine it with LAYOUT mode, as below:
SELECT AI_PARSE_DOCUMENT (
TO_FILE('@docs.doc_stage','research-paper-example.pdf'),
{'mode': 'LAYOUT' , 'page_split': true}) AS research_paper_example;There is one restriction to remember: page_split works only with paginated formats. PDF, PowerPoint (.pptx) and Word (.docx) are supported. Any other format returns an error, even though AI_PARSE_DOCUMENT can parse those formats without splitting.
Checkpoint 4 of 8· Check yourself
A pipeline sets 'page_split': TRUE for every staged file. Which file will cause an error?
page_split supports only PDF, PPTX and DOCX. An image such as a PNG returns an error when page_split is set.
“This feature supports only PDF, PowerPoint (.pptx), and Word (.docx) documents. Documents in other formats return an error.”Source: docs.snowflake.com
Checkpoint 5 of 8· Exam question
A finance analyst uploads a multi-page quarterly report PDF that contains several data tables and needs the parsed output to preserve table structure and section headings for downstream processing. Which AI_PARSE_DOCUMENT configuration best satisfies this requirement?
Correct answer: A — Set mode to LAYOUT, so the response markdown preserves tables and structural headings alongside the extracted text
- A. LAYOUT mode extracts layout as well as text, including structural content such as tables, which is exactly what preserving table and heading structure requires.
- B. OCR mode extracts text only and does not preserve layout, so it cannot reconstruct table borders or column alignment as this option claims.
- C. Splitting a document into individual pages changes how output is chunked but does not add structural table or heading preservation, which only LAYOUT mode provides.
- D. Extract_images pulls embedded picture content and requires LAYOUT mode to function, but it does not convert tabular text into structured output on its own.
Sources1
4.Limiting pages: page_filter and the per-document ceiling
The exam guide lists a 'page_limit' topic, but the documentation available here defines no option with that name. Two documented mechanisms limit which pages are processed. The first is the 'page_filter' option, which you control. The second is a fixed limit on document size: 100 MB per file and 2,000 pages per document.
'page_filter' takes an array of ranges. Each range is an object with 'start' and 'end' fields, page indexes start at 0, start is inclusive and end is exclusive. So {'start': 0, 'end': 1} selects only the first page. Setting page_filter implies page_split, so you do not need to set both, and the result comes back as a "pages" array. The documentation suggests this when you already know which pages hold the information you need.
SELECT AI_PARSE_DOCUMENT(
TO_FILE('@my_documents', 'ResearchArticle.pdf'),
{'mode': 'LAYOUT', 'page_filter': [{'start': 0, 'end': 1}]} );Checkpoint 6 of 8· Fill the gap
Which option key restricts parsing to the first page here?
SELECT AI_PARSE_DOCUMENT(
TO_FILE('@my_documents', 'ResearchArticle.pdf'),
{'mode': 'LAYOUT', ' ? ': [{'start': 0, 'end': 1}]} );page_filter takes an array of start/end ranges. page_split only turns splitting on, and page_limit and page_range are not documented option keys.
Source: docs.snowflake.comCheckpoint 7 of 8· Check yourself
Which page_filter returns the third and fourth pages of a PDF?
Pages are 0-indexed, so the third and fourth pages are indexes 2 and 3. Because end is exclusive, the range has to end at 4.
“Each range is an object with start and end fields that specify the first (inclusive) and last (exclusive) page in the range.”Source: docs.snowflake.com
Sources3
5.When parsing fails
By default, AI_PARSE_DOCUMENT returns NULL for a document it cannot process. In a multi-row query, the failing rows return NULL and the other rows still complete. To find out why a row failed, pass TRUE as the third argument, return_error_details. The function then returns an OBJECT with three top-level fields: value holds the parsed data, error holds the message (or NULL on success), and metadata holds document metadata such as pageCount.
Checkpoint 8 of 8· Check yourself
A query parses 5,000 staged PDFs without return_error_details, and three files are corrupt. What happens?
Without return_error_details, a failed row returns NULL and does not stop the query. You only get the error OBJECT when you pass TRUE.
“rows with errors return NULL and don’t prevent the query from completing.”Source: docs.snowflake.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Because the documentation recommends LAYOUT for most use cases, AI_PARSE_DOCUMENT uses LAYOUT unless told otherwise.Why is that wrong?
OCR is the default. To get tables and headings back as Markdown, you have to set 'mode': 'LAYOUT' explicitly.
Covered in OCR mode vs LAYOUT mode
2.'extract_images': TRUE works in any mode.Why is that wrong?
Image extraction is available only in LAYOUT mode.
Covered in OCR mode vs LAYOUT mode
3.To use page_filter you must also set 'page_split': TRUE.Why is that wrong?
page_filter implies page_split, so setting both is redundant.
Covered in Limiting pages: page_filter and the per-document ceiling
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A FILE object that specifies the document to parse, stored in a Snowflake stage.”
↩︎ The shape of an AI_PARSE_DOCUMENT call“'LAYOUT': The function extracts layout as well as text, including structural content such as tables.”
↩︎ OCR mode vs LAYOUT mode“To process long documents that exceed the token limit of AI_PARSE_DOCUMENT, set this option to TRUE.”
↩︎ page_split: one result per page“The page index in the file, starting at 0. Page numbers and formats specified in the document are ignored.”
↩︎ page_split: one result per page“By default, if AI_PARSE_DOCUMENT can’t process the input, the function returns NULL.”
↩︎ When parsing fails“Returns the extracted content from a document on a Snowflake stage as a JSON-formatted string.”
↩︎ Key concept“'OCR': The function extracts text only. This is the default mode.”
↩︎ Exam trap 1“'extract_images': If set to TRUE, the function extracts images embedded in the document. Requires LAYOUT mode.”
↩︎ Exam trap 2“Specifying page_filter implies page_split.”
↩︎ Exam trap 3“'OCR': The function extracts text only. This is the default mode.”
↩︎ Prediction“Plain text (in OCR mode) or Markdown-formatted text (in LAYOUT mode).”
↩︎ Checkpoint“This feature supports only PDF, PowerPoint (.pptx), and Word (.docx) documents. Documents in other formats return an error.”
↩︎ Checkpoint“Each range is an object with start and end fields that specify the first (inclusive) and last (exclusive) page in the range.”
↩︎ Checkpoint“rows with errors return NULL and don’t prevent the query from completing.”
↩︎ Checkpoint - 2.
“This legacy function will be deprecated by the end of 2026.”
↩︎ The shape of an AI_PARSE_DOCUMENT call - 3.
“LAYOUT mode is the preferred choice for most use cases, especially for complex documents.”
↩︎ OCR mode vs LAYOUT mode“OCR mode is recommended for fast, high-quality text extraction from scanned or text-heavy documents such as contracts, insurance claims, and manuals.”
↩︎ OCR mode vs LAYOUT mode“There is no additional cost for using this parameter.”
↩︎ OCR mode vs LAYOUT mode“Maximum pages per document | 2,000”
↩︎ Limiting pages: page_filter and the per-document ceiling“This is useful when you know what pages the information you’re looking for is on.”
↩︎ Limiting pages: page_filter and the per-document ceiling