CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 4 · Lesson 12/15

    AI_PARSE_DOCUMENT: OCR vs LAYOUT, page_split and page limits

    Use document parsing functions.

    11 min read
    3.75% of exam
    3 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Write an AI_PARSE_DOCUMENT call against a staged file using TO_FILE and an options object
    • Choose between OCR and LAYOUT mode for a given document and know what each returns
    • Predict how page_split changes the shape of the JSON output and which file formats it accepts
    • Restrict parsing to specific pages with page_filter and state the per-document page ceiling
    • Describe what AI_PARSE_DOCUMENT returns when a document cannot be parsed

    Key concept

    Parsing mode decides the output shape — AI_PARSE_DOCUMENT turns a staged document into a JSON string. The options object controls what comes back: OCR mode returns plain text, LAYOUT mode returns Markdown that keeps the document's structure, and page options decide whether you get one content field or a pages array.

    1.The shape of an AI_PARSE_DOCUMENT call

    AI_PARSE_DOCUMENT is the Cortex AI function that turns a document into text Snowflake can work with. It reads files on internal or external stages and keeps the reading order and structural elements such as tables and headers. Its output can then feed other Cortex AI functions for extraction, classification or enrichment. The function takes one required argument and two optional ones:

    AI_PARSE_DOCUMENT signaturesql
    AI_PARSE_DOCUMENT( <file_object> [, <options> ] [, <return_error_details> ] )

    The first argument is a FILE object pointing to a document on a Snowflake stage, normally built with TO_FILE('@stage', 'path'). The second is an OBJECT of options. Every key in it is optional, and the keys you will meet most are 'mode', 'page_split', 'page_filter' and 'extract_images'. The function returns a JSON-formatted string, so the documentation suggests passing it through PARSE_JSON before using the result in SQL.

    The older function, SNOWFLAKE.CORTEX.PARSE_DOCUMENT, was called differently: it took a stage name and a relative path as two separate string arguments. Its reference page is kept only for backward compatibility and points new work to AI_PARSE_DOCUMENT.

    Checkpoint 1 of 8· Check yourself

    You are writing new code to parse a PDF staged at @docs.doc_stage. Which call matches the current function?

    Sources12

    2.OCR mode vs LAYOUT mode

    The 'mode' key picks one of two extraction types. OCR extracts text only and is the default. LAYOUT extracts layout as well as text, including structural content such as tables. The two modes also return different formats: content is plain text in OCR mode and Markdown in LAYOUT mode. Headings come back as # lines, and tables come back as Markdown tables.

    The defaults and the recommendations point in different directions. The user guide calls LAYOUT the preferred choice for most use cases, especially complex documents, and it is required for image extraction. It recommends OCR for fast, high-quality text extraction from scanned or text-heavy documents such as contracts, insurance claims and manuals. So if you want tables or headings to survive, you have to ask for LAYOUT explicitly.

    The two parsing modes side by side
    AspectOCRLAYOUT
    What is extractedText onlyText plus layout, including tables and headers
    Format of contentPlain textMarkdown-formatted text
    Default?YesNo: set 'mode': 'LAYOUT'
    Recommended forFast text extraction from scanned or text-heavy documentsMost use cases, especially complex documents
    'extract_images'Not availableAvailable, at no additional cost
    Document Processing Playground tabTextMarkdown

    Image extraction is only available in LAYOUT. Setting 'extract_images' to TRUE adds an "images" array to the output. Each entry has an id, bounding-box coordinates (top_left_x, top_left_y, bottom_right_x, bottom_right_y) and the image itself as image_base64. The extracted images can then go to AI_EXTRACT or AI_COMPLETE for tagging and analysis.

    Checkpoint 2 of 8· Match them up

    Match each option or mode to what it produces

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 8· Exam question

    A claims processing team receives scanned insurance claim PDFs that are dense with body text and have almost no tables. They need AI_PARSE_DOCUMENT to extract the text quickly and cost-effectively for downstream analysis. Which configuration best fits this requirement?

    Sources31

    3.page_split: one result per page

    By default, AI_PARSE_DOCUMENT returns the whole document as a single "content" field. When 'page_split' is TRUE, the document is split into pages and each page is processed separately. The output then becomes a "pages" array, and each element has a "content" field and an "index" field. The index starts at 0, and any page numbering printed in the document is ignored. A one-page document still comes back as a "pages" array with one element, so downstream code can rely on the same shape every time.

    This is the documented way to handle long documents. The reference tip says to set the option to TRUE for documents that exceed the function's token limit. The user guide's examples combine it with LAYOUT mode, as below:

    LAYOUT mode with page_split on a two-column research papersql
    SELECT AI_PARSE_DOCUMENT (
        TO_FILE('@docs.doc_stage','research-paper-example.pdf'),
        {'mode': 'LAYOUT' , 'page_split': true}) AS research_paper_example;

    There is one restriction to remember: page_split works only with paginated formats. PDF, PowerPoint (.pptx) and Word (.docx) are supported. Any other format returns an error, even though AI_PARSE_DOCUMENT can parse those formats without splitting.

    Checkpoint 4 of 8· Check yourself

    A pipeline sets 'page_split': TRUE for every staged file. Which file will cause an error?

    Checkpoint 5 of 8· Exam question

    A finance analyst uploads a multi-page quarterly report PDF that contains several data tables and needs the parsed output to preserve table structure and section headings for downstream processing. Which AI_PARSE_DOCUMENT configuration best satisfies this requirement?

    Sources1

    4.Limiting pages: page_filter and the per-document ceiling

    The exam guide lists a 'page_limit' topic, but the documentation available here defines no option with that name. Two documented mechanisms limit which pages are processed. The first is the 'page_filter' option, which you control. The second is a fixed limit on document size: 100 MB per file and 2,000 pages per document.

    'page_filter' takes an array of ranges. Each range is an object with 'start' and 'end' fields, page indexes start at 0, start is inclusive and end is exclusive. So {'start': 0, 'end': 1} selects only the first page. Setting page_filter implies page_split, so you do not need to set both, and the result comes back as a "pages" array. The documentation suggests this when you already know which pages hold the information you need.

    Parsing only the first page of a 55-page research papersql
    SELECT AI_PARSE_DOCUMENT(
      TO_FILE('@my_documents', 'ResearchArticle.pdf'),
      {'mode': 'LAYOUT', 'page_filter': [{'start': 0, 'end': 1}]} );

    Checkpoint 6 of 8· Fill the gap

    Which option key restricts parsing to the first page here?

    SELECT AI_PARSE_DOCUMENT(
      TO_FILE('@my_documents', 'ResearchArticle.pdf'),
      {'mode': 'LAYOUT', ' ? ': [{'start': 0, 'end': 1}]} );

    Checkpoint 7 of 8· Check yourself

    Which page_filter returns the third and fourth pages of a PDF?

    Sources3

    5.When parsing fails

    By default, AI_PARSE_DOCUMENT returns NULL for a document it cannot process. In a multi-row query, the failing rows return NULL and the other rows still complete. To find out why a row failed, pass TRUE as the third argument, return_error_details. The function then returns an OBJECT with three top-level fields: value holds the parsed data, error holds the message (or NULL on success), and metadata holds document metadata such as pageCount.

    Checkpoint 8 of 8· Check yourself

    A query parses 5,000 staged PDFs without return_error_details, and three files are corrupt. What happens?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Because the documentation recommends LAYOUT for most use cases, AI_PARSE_DOCUMENT uses LAYOUT unless told otherwise.Why is that wrong?

      OCR is the default. To get tables and headings back as Markdown, you have to set 'mode': 'LAYOUT' explicitly.

      Covered in OCR mode vs LAYOUT mode

    2. 2.'extract_images': TRUE works in any mode.Why is that wrong?

      Image extraction is available only in LAYOUT mode.

      Covered in OCR mode vs LAYOUT mode

    3. 3.To use page_filter you must also set 'page_split': TRUE.Why is that wrong?

      page_filter implies page_split, so setting both is redundant.

      Covered in Limiting pages: page_filter and the per-document ceiling

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “A FILE object that specifies the document to parse, stored in a Snowflake stage.”
      ↩︎ The shape of an AI_PARSE_DOCUMENT call
      “'LAYOUT': The function extracts layout as well as text, including structural content such as tables.”
      ↩︎ OCR mode vs LAYOUT mode
      “To process long documents that exceed the token limit of AI_PARSE_DOCUMENT, set this option to TRUE.”
      ↩︎ page_split: one result per page
      “The page index in the file, starting at 0. Page numbers and formats specified in the document are ignored.”
      ↩︎ page_split: one result per page
      “By default, if AI_PARSE_DOCUMENT can’t process the input, the function returns NULL.”
      ↩︎ When parsing fails
      “Returns the extracted content from a document on a Snowflake stage as a JSON-formatted string.”
      ↩︎ Key concept
      “'OCR': The function extracts text only. This is the default mode.”
      ↩︎ Exam trap 1
      “'extract_images': If set to TRUE, the function extracts images embedded in the document. Requires LAYOUT mode.”
      ↩︎ Exam trap 2
      “Specifying page_filter implies page_split.”
      ↩︎ Exam trap 3
      “'OCR': The function extracts text only. This is the default mode.”
      ↩︎ Prediction
      “Plain text (in OCR mode) or Markdown-formatted text (in LAYOUT mode).”
      ↩︎ Checkpoint
      “This feature supports only PDF, PowerPoint (.pptx), and Word (.docx) documents. Documents in other formats return an error.”
      ↩︎ Checkpoint
      “Each range is an object with start and end fields that specify the first (inclusive) and last (exclusive) page in the range.”
      ↩︎ Checkpoint
      “rows with errors return NULL and don’t prevent the query from completing.”
      ↩︎ Checkpoint
    2. 3.
      “LAYOUT mode is the preferred choice for most use cases, especially for complex documents.”
      ↩︎ OCR mode vs LAYOUT mode
      “OCR mode is recommended for fast, high-quality text extraction from scanned or text-heavy documents such as contracts, insurance claims, and manuals.”
      ↩︎ OCR mode vs LAYOUT mode
      “There is no additional cost for using this parameter.”
      ↩︎ OCR mode vs LAYOUT mode
      “This is useful when you know what pages the information you’re looking for is on.”
      ↩︎ Limiting pages: page_filter and the per-document ceiling

    Continue to page 2 of 2

    AI_EXTRACT: response formats and prompting for document extraction

    Spotted a mistake, or was something unclear? Tell us.