CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 2 · Lesson 4/15

    Unstructured Data Analysis in Snowflake: AI_PARSE_DOCUMENT, AI_EXTRACT, Cortex Search and Multimodal AI

    Perform data analysis given a use case.

    12 min read
    7.6% of exam
    6 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Choose between LAYOUT and OCR mode of AI_PARSE_DOCUMENT for a given document set
    • Design AI_EXTRACT entity, list and table extractions that stay accurate and within the output limit
    • Chunk text with SPLIT_TEXT_RECURSIVE_CHARACTER and size chunks for Cortex Search
    • Pick a Cortex Search embedding model and decide whether to keep semantic reranking
    • Match audio, image and similarity tasks to AI_TRANSCRIBE, AI_COMPLETE, AI_MULTI_EMBED and AI_SIMILARITY

    Key concept

    Retrieval augmented generation (RAG) — Before the LLM answers, you fetch the relevant pieces of your own content and give them to the model, so the answer rests on your data and not only on what the model learned in training. Parsing, extraction, chunking and Cortex Search are the steps that get unstructured content into a form that can be retrieved this way.

    1.Turning files into text: AI_PARSE_DOCUMENT

    Most unstructured analysis starts with files such as PDFs, scans and filings sitting on a stage. AI_PARSE_DOCUMENT reads documents from internal or external stages and keeps the reading order and structural elements such as tables and headers. You can then pass its output to other Cortex AI functions for extraction, classification or enrichment, or index it for RAG.

    The exam tests the difference between its two modes. LAYOUT mode is the default recommendation, especially for complex documents. It is optimized for structured content such as tables, headers and layout relationships, and it is the only mode that can extract images (set 'extract_images': true, which costs nothing extra). OCR mode is the faster choice when you only need high-quality text from scanned or text-heavy documents such as contracts, insurance claims and manuals.

    LAYOUT mode with page_split on a two-column research paper; the response is JSON with Markdown content per pagesql
    SELECT AI_PARSE_DOCUMENT (
        TO_FILE('@docs.doc_stage','research-paper-example.pdf'),
        {'mode': 'LAYOUT' , 'page_split': true}) AS research_paper_example;

    In both modes, page_split returns each page of a multi-page document separately, and page_filter limits processing to the pages you list. Setting page_filter turns on page_split automatically. LAYOUT output is Markdown, so headers come back as # lines and tables as Markdown tables. That matters later, because a Markdown-aware splitter can use that structure when it chunks the text.

    Checkpoint 1 of 7· Check yourself

    A team needs to index native PDFs and keep their section headers and tables, and also wants to pull out embedded charts for later analysis. Which configuration fits?

    Checkpoint 2 of 7· Exam question

    A team has a stage full of scanned insurance claim PDFs (image-only, no embedded text layer) and needs the full document content converted into text that preserves reading order and section layout for downstream semantic search. Which function and mode should they use?

    Sources1

    2.Pulling structured fields out: AI_EXTRACT

    Use AI_EXTRACT when you want specific values out of a document, not the whole text. It runs on arctic-extract, a vision-based LLM, and it can read graphical content as well as text-heavy paragraphs, including logos, handwritten text such as signatures, tables and checkmarks. It returns three shapes of answer:

    - Entity: a natural-language question or description, for example the city or the ZIP code. - List: a JSON schema for an array, for example every account holder on a bank statement. - Table: a JSON schema that gives the table title and the columns to extract.

    Questions should be in plain English, specific, and ask for one value each. "What is the date?" is a poor question for a document that contains both an issue date and a signature date.

    Most table-extraction accuracy comes from how you define the columns. List them in the order they appear in the document, and put repeated values such as Invoice Number first. Copy column names exactly as the document writes them (Product Code, not product_code). Add a Section column when the table is split into named sections. For hierarchical headers, join the parent names into the column name. The output limit is the constraint most people miss: table answers stop at 4096 tokens. For a table that runs across several pages, split the document into one-page documents and join the results afterwards. If even one page is too dense, split the table by columns.

    Checkpoint 3 of 7· Check yourself

    An AI_EXTRACT table extraction over a 12-page transaction table returns rows only from the first few pages. What is the documented fix?

    Sources2

    3.Preparing text for retrieval: recursive and Markdown splitting

    For search and RAG, long parsed text has to be cut into chunks. SNOWFLAKE.CORTEX.SPLIT_TEXT_RECURSIVE_CHARACTER first splits on the highest-priority separator. Any chunk that is still longer than chunk_size is split again on the next separator, and this repeats until every chunk fits. The format argument sets the default separators. With 'none', only the separator list is used, and the default list is paragraph break, line break, space, then any character. With 'markdown', the text is also split on headers, code blocks and tables, so it works well on AI_PARSE_DOCUMENT LAYOUT output. The optional overlap repeats some characters from the previous chunk so each chunk keeps some context.

    Checkpoint 4 of 7· Fill the gap

    This query chunks Markdown documents and keeps their header, code-block and table boundaries. Which format value completes it?

    SELECT
       doc_id,
       c.value
    FROM
       sample_documents,
       LATERAL FLATTEN( input => SNOWFLAKE.CORTEX.SPLIT_TEXT_RECURSIVE_CHARACTER (
          document,
           ? ,
          25,
          10
       )) c;

    Chunk sizing is one of the few retrieval-quality settings you control yourself. Even though models such as snowflake-arctic-embed-l-v2.0-8k accept long inputs, smaller chunks usually give better retrieval and better downstream LLM answers. Two details lead to wrong answers on the exam. First, chunk_size in the splitter is counted in characters, while the 512 recommendation is in tokens, at roughly 4 characters per token. Second, text longer than the embedding model's context window is not rejected. Cortex Search truncates it before embedding, but keyword retrieval still uses the full text.

    Checkpoint 5 of 7· Check yourself

    A search column holds 3,000-token values and the service uses an embedding model with a 512-token context window. What happens?

    Sources34

    Cortex Search is the managed retrieval layer for RAG. Every query combines vector search for semantic matches, keyword search for lexical matches, and a semantic reranking step that reorders the top results. You choose the embedding model with EMBEDDING_MODEL when you create the service. Embedding, indexing and refresh are managed for you, and TARGET_LAG controls how often the service picks up changes.

    Creating a service with an explicit multilingual embedding modelsql
    CREATE OR REPLACE CORTEX SEARCH SERVICE transcript_search_service
      ON transcript_text
      ATTRIBUTES region
      WAREHOUSE = cortex_search_wh
      TARGET_LAG = '1 day'
      EMBEDDING_MODEL = 'snowflake-arctic-embed-l-v2.0'
      AS (
        SELECT
            transcript_text,
            region,
            agent_id
        FROM support_transcripts
    );
    Embedding models available in Cortex Search
    ModelDimensionsContext window (tokens)Languages
    snowflake-arctic-embed-m-v1.5 (default)768512English-only; fastest indexing
    snowflake-arctic-embed-l-v2.01024512Multilingual
    snowflake-arctic-embed-l-v2.0-8k10248192Multilingual
    voyage-multilingual-2102432,000Multilingual

    Reranking is on by default and trades latency for quality. Turning it off with "reranker": "none" in scoring_config saves 100–300 ms per query on average. How much quality you lose depends on the workload, so Snowflake advises comparing results with and without reranking before you disable it. You can also change the relative weights of the text, vector and reranker scores. Reranking does not apply to batch search queries.

    Disabling the reranker for a single queryjson
    {
      "scoring_config": {
          "reranker": "none"
      }
    }

    Checkpoint 6 of 7· Check yourself

    A search-bar application needs faster responses and can accept a small drop in relevance. Which change is designed for this trade-off?

    Sources45

    5.Multimodal analytics: audio, images and similarity

    Unstructured data also includes audio and images, and the same AI functions handle them. AI_COMPLETE takes images as well as text, and TO_FILE creates the stage-file reference that AI_COMPLETE and other file-accepting functions need. AI_TRANSCRIBE converts staged audio and video into text with timestamps and speaker information, and from there the text-analysis functions take over. AI_PARSE_DOCUMENT with extract_images gives you a document's embedded images, which you can then tag or analyze with AI_EXTRACT or AI_COMPLETE. AI_PARSE_DOCUMENT names multimodal RAG as one of its use cases.

    For comparing content, AI_EMBED produces an embedding vector for text or an image. AI_MULTI_EMBED creates embeddings from text, images, audio or video, with segment metadata, for semantic search across modalities. AI_SIMILARITY calculates the embedding similarity between two inputs, which is useful for questions like "how close is this ticket to that known issue". The sources here give only that one-line description of AI_SIMILARITY. They do not list its supported input types or options.

    Checkpoint 7 of 7· Match them up

    Match each task to the function built for it

    Tap a term, then the definition that fits it.

    Sources61

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.OCR mode is the right choice whenever documents contain tables, because it is the higher-quality mode.Why is that wrong?

      OCR mode is for fast text extraction from scanned or text-heavy documents. Tables, headers, layout relationships and image extraction all need LAYOUT mode.

      Covered in Turning files into text: AI_PARSE_DOCUMENT

    2. 2.Setting chunk_size to 512 in SPLIT_TEXT_RECURSIVE_CHARACTER gives you the recommended 512-token chunks.Why is that wrong?

      chunk_size counts characters, not tokens. Since a token is about 4 characters, a 512-token chunk is roughly 2,000 characters.

      Covered in Preparing text for retrieval: recursive and Markdown splitting

    3. 3.Picking the longest-context embedding model and indexing whole documents gives the best retrieval quality.Why is that wrong?

      Smaller chunks usually give better retrieval and better downstream LLM answers, even when a long-context model is available.

      Covered in Preparing text for retrieval: recursive and Markdown splitting

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “OCR mode is recommended for fast, high-quality text extraction from scanned or text-heavy documents such as contracts, insurance claims, and manuals.”
      ↩︎ Turning files into text: AI_PARSE_DOCUMENT
      “If using page_filter, page_split is implied, and you do not need to set it explicitly.”
      ↩︎ Turning files into text: AI_PARSE_DOCUMENT
      “AI_PARSE_DOCUMENT returns the content in Markdown format.”
      ↩︎ Turning files into text: AI_PARSE_DOCUMENT
      “Image understanding: Use extracted images with AI_EXTRACT or AI_COMPLETE for automatic tagging and analysis”
      ↩︎ Multimodal analytics: audio, images and similarity
      “LAYOUT mode is the preferred choice for most use cases, especially for complex documents.”
      ↩︎ Exam trap 1
      “It is optimized for extracting structured content such as tables, headers, and layout relationships, and is required for image extraction.”
      ↩︎ Checkpoint
    2. 2.
      “AI_EXTRACT uses arctic-extract, a proprietary vision-based large language model (LLM) that delivers high extraction accuracy.”
      ↩︎ Pulling structured fields out: AI_EXTRACT
      “content in a graphical form, such as logos, handwritten text (for example, signatures), tables, or checkmarks”
      ↩︎ Pulling structured fields out: AI_EXTRACT
      “Ask for a single value in each question.”
      ↩︎ Pulling structured fields out: AI_EXTRACT
      “The model for table extraction returns answers that are up to 4096 tokens long.”
      ↩︎ Pulling structured fields out: AI_EXTRACT
      “If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”
      ↩︎ Checkpoint
    3. 3.
      “markdown: Separates on headers, code blocks, and tables, in addition to any separators in the separators field.”
      ↩︎ Preparing text for retrieval: recursive and Markdown splitting
      “Overlap is useful for ensuring that each chunk has some context about the previous chunk.”
      ↩︎ Preparing text for retrieval: recursive and Markdown splitting
      “An integer specifying the maximum number of characters in each chunk.”
      ↩︎ Exam trap 2
    4. 4.
      “Cortex Search truncates the string to the size of the context window before embedding it into vector space for semantic search.”
      ↩︎ Preparing text for retrieval: recursive and Markdown splitting
      “Semantic reranking for reranking the most relevant documents in the result set.”
      ↩︎ Cortex Search: embedding models and semantic reranking
      “Retrieval augmented generation (RAG) is a technique for retrieving data from a knowledge base to enhance the generated response of a large language model.”
      ↩︎ Key concept
      “research shows that a smaller chunk size typically results in higher retrieval and downstream LLM response quality”
      ↩︎ Exam trap 3
      “Snowflake recommends splitting the text in your search column into chunks of no more than 512 tokens (about 385 English words).”
      ↩︎ Prediction
      “However, Cortex Search uses the full body of text for keyword-based retrieval.”
      ↩︎ Checkpoint
    5. 5.
      “Disabling reranking reduces query latency by 100-300ms on average”
      ↩︎ Cortex Search: embedding models and semantic reranking
      “Reranking is not supported for batch search queries.”
      ↩︎ Cortex Search: embedding models and semantic reranking
      “While reranking can measurably increase result relevance, it can also noticeably increase query latency.”
      ↩︎ Checkpoint
    6. 6.
      “AI_COMPLETE: Generates a completion for a given text string or image using a selected LLM.”
      ↩︎ Multimodal analytics: audio, images and similarity
      “AI_SIMILARITY: Calculates the embedding similarity between two inputs.”
      ↩︎ Multimodal analytics: audio, images and similarity
      “AI_MULTI_EMBED: Creates multimodal embeddings from text, images, audio, or video, returning one or more vectors with segment metadata for semantic search across modalities.”
      ↩︎ Multimodal analytics: audio, images and similarity
      “AI_TRANSCRIBE: Transcribes audio and video files stored in a stage, extracting text, timestamps, and speaker information.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Structured Data Analysis and Model Choice: Cortex Analyst, Verified Queries and Provisioned Throughput

    Spotted a mistake, or was something unclear? Tell us.