What you will be able to do
- Choose between LAYOUT and OCR mode of AI_PARSE_DOCUMENT for a given document set
- Design AI_EXTRACT entity, list and table extractions that stay accurate and within the output limit
- Chunk text with SPLIT_TEXT_RECURSIVE_CHARACTER and size chunks for Cortex Search
- Pick a Cortex Search embedding model and decide whether to keep semantic reranking
- Match audio, image and similarity tasks to AI_TRANSCRIBE, AI_COMPLETE, AI_MULTI_EMBED and AI_SIMILARITY
Key concept
Retrieval augmented generation (RAG) — Before the LLM answers, you fetch the relevant pieces of your own content and give them to the model, so the answer rests on your data and not only on what the model learned in training. Parsing, extraction, chunking and Cortex Search are the steps that get unstructured content into a form that can be retrieved this way.
1.Turning files into text: AI_PARSE_DOCUMENT
Most unstructured analysis starts with files such as PDFs, scans and filings sitting on a stage. AI_PARSE_DOCUMENT reads documents from internal or external stages and keeps the reading order and structural elements such as tables and headers. You can then pass its output to other Cortex AI functions for extraction, classification or enrichment, or index it for RAG.
The exam tests the difference between its two modes. LAYOUT mode is the default recommendation, especially for complex documents. It is optimized for structured content such as tables, headers and layout relationships, and it is the only mode that can extract images (set 'extract_images': true, which costs nothing extra). OCR mode is the faster choice when you only need high-quality text from scanned or text-heavy documents such as contracts, insurance claims and manuals.
SELECT AI_PARSE_DOCUMENT (
TO_FILE('@docs.doc_stage','research-paper-example.pdf'),
{'mode': 'LAYOUT' , 'page_split': true}) AS research_paper_example;In both modes, page_split returns each page of a multi-page document separately, and page_filter limits processing to the pages you list. Setting page_filter turns on page_split automatically. LAYOUT output is Markdown, so headers come back as # lines and tables as Markdown tables. That matters later, because a Markdown-aware splitter can use that structure when it chunks the text.
Checkpoint 1 of 7· Check yourself
A team needs to index native PDFs and keep their section headers and tables, and also wants to pull out embedded charts for later analysis. Which configuration fits?
LAYOUT mode is the one optimized for tables, headers and layout relationships, and only LAYOUT mode can extract images. page_filter selects which pages to process. It has nothing to do with images.
“It is optimized for extracting structured content such as tables, headers, and layout relationships, and is required for image extraction.”Source: docs.snowflake.com
Checkpoint 2 of 7· Exam question
A team has a stage full of scanned insurance claim PDFs (image-only, no embedded text layer) and needs the full document content converted into text that preserves reading order and section layout for downstream semantic search. Which function and mode should they use?
Correct answer: A — AI_PARSE_DOCUMENT in LAYOUT mode
- A. AI_PARSE_DOCUMENT in LAYOUT mode is built to convert scanned or digital-native documents into structured, reading-order text while preserving layout such as headings, tables, and sections, which is exactly what downstream chunking and semantic search need.
- B. AI_EXTRACT is designed to pull specific fields into a defined schema rather than produce a full-text, layout-preserving representation of the whole document, so it is not the right first step for general semantic search.
- C. Feeding raw scanned pages directly to a vision model with no dedicated parsing step skips the structured layout extraction and page/table handling that AI_PARSE_DOCUMENT provides, making it a less reliable and less scalable approach.
- D. AI_CLASSIFY assigns categorical labels to documents or pages; it does not convert scanned content into full text and cannot preserve document structure.
Sources1
2.Pulling structured fields out: AI_EXTRACT
Use AI_EXTRACT when you want specific values out of a document, not the whole text. It runs on arctic-extract, a vision-based LLM, and it can read graphical content as well as text-heavy paragraphs, including logos, handwritten text such as signatures, tables and checkmarks. It returns three shapes of answer:
- Entity: a natural-language question or description, for example the city or the ZIP code. - List: a JSON schema for an array, for example every account holder on a bank statement. - Table: a JSON schema that gives the table title and the columns to extract.
Questions should be in plain English, specific, and ask for one value each. "What is the date?" is a poor question for a document that contains both an issue date and a signature date.
Most table-extraction accuracy comes from how you define the columns. List them in the order they appear in the document, and put repeated values such as Invoice Number first. Copy column names exactly as the document writes them (Product Code, not product_code). Add a Section column when the table is split into named sections. For hierarchical headers, join the parent names into the column name. The output limit is the constraint most people miss: table answers stop at 4096 tokens. For a table that runs across several pages, split the document into one-page documents and join the results afterwards. If even one page is too dense, split the table by columns.
Checkpoint 3 of 7· Check yourself
An AI_EXTRACT table extraction over a 12-page transaction table returns rows only from the first few pages. What is the documented fix?
Table answers are capped at 4096 tokens, so extraction stops once it reaches the cap. Extracting page by page keeps each answer under the cap. Renaming columns to snake_case goes against the guidance to copy the document's own names.
“If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”Source: docs.snowflake.com
Sources2
3.Preparing text for retrieval: recursive and Markdown splitting
For search and RAG, long parsed text has to be cut into chunks. SNOWFLAKE.CORTEX.SPLIT_TEXT_RECURSIVE_CHARACTER first splits on the highest-priority separator. Any chunk that is still longer than chunk_size is split again on the next separator, and this repeats until every chunk fits. The format argument sets the default separators. With 'none', only the separator list is used, and the default list is paragraph break, line break, space, then any character. With 'markdown', the text is also split on headers, code blocks and tables, so it works well on AI_PARSE_DOCUMENT LAYOUT output. The optional overlap repeats some characters from the previous chunk so each chunk keeps some context.
Checkpoint 4 of 7· Fill the gap
This query chunks Markdown documents and keeps their header, code-block and table boundaries. Which format value completes it?
SELECT
doc_id,
c.value
FROM
sample_documents,
LATERAL FLATTEN( input => SNOWFLAKE.CORTEX.SPLIT_TEXT_RECURSIVE_CHARACTER (
document,
? ,
25,
10
)) c;'markdown' adds header, code-block and table separators to the defaults. 'none' uses only the separators list, and 'layout' is a mode of AI_PARSE_DOCUMENT, not a format of this function.
Source: docs.snowflake.comChunk sizing is one of the few retrieval-quality settings you control yourself. Even though models such as snowflake-arctic-embed-l-v2.0-8k accept long inputs, smaller chunks usually give better retrieval and better downstream LLM answers. Two details lead to wrong answers on the exam. First, chunk_size in the splitter is counted in characters, while the 512 recommendation is in tokens, at roughly 4 characters per token. Second, text longer than the embedding model's context window is not rejected. Cortex Search truncates it before embedding, but keyword retrieval still uses the full text.
Checkpoint 5 of 7· Check yourself
A search column holds 3,000-token values and the service uses an embedding model with a 512-token context window. What happens?
Values that are too long are truncated for semantic embedding only. Nothing splits them automatically, which is why you chunk them yourself beforehand.
“However, Cortex Search uses the full body of text for keyword-based retrieval.”Source: docs.snowflake.com
4.Cortex Search: embedding models and semantic reranking
Cortex Search is the managed retrieval layer for RAG. Every query combines vector search for semantic matches, keyword search for lexical matches, and a semantic reranking step that reorders the top results. You choose the embedding model with EMBEDDING_MODEL when you create the service. Embedding, indexing and refresh are managed for you, and TARGET_LAG controls how often the service picks up changes.
CREATE OR REPLACE CORTEX SEARCH SERVICE transcript_search_service
ON transcript_text
ATTRIBUTES region
WAREHOUSE = cortex_search_wh
TARGET_LAG = '1 day'
EMBEDDING_MODEL = 'snowflake-arctic-embed-l-v2.0'
AS (
SELECT
transcript_text,
region,
agent_id
FROM support_transcripts
);| Model | Dimensions | Context window (tokens) | Languages |
|---|---|---|---|
| snowflake-arctic-embed-m-v1.5 (default) | 768 | 512 | English-only; fastest indexing |
| snowflake-arctic-embed-l-v2.0 | 1024 | 512 | Multilingual |
| snowflake-arctic-embed-l-v2.0-8k | 1024 | 8192 | Multilingual |
| voyage-multilingual-2 | 1024 | 32,000 | Multilingual |
Reranking is on by default and trades latency for quality. Turning it off with "reranker": "none" in scoring_config saves 100–300 ms per query on average. How much quality you lose depends on the workload, so Snowflake advises comparing results with and without reranking before you disable it. You can also change the relative weights of the text, vector and reranker scores. Reranking does not apply to batch search queries.
{
"scoring_config": {
"reranker": "none"
}
}Checkpoint 6 of 7· Check yourself
A search-bar application needs faster responses and can accept a small drop in relevance. Which change is designed for this trade-off?
Reranking improves relevance but adds latency, and disabling it per query is the documented way to trade quality for speed. Batch search ignores reranker settings altogether.
“While reranking can measurably increase result relevance, it can also noticeably increase query latency.”Source: docs.snowflake.com
5.Multimodal analytics: audio, images and similarity
Unstructured data also includes audio and images, and the same AI functions handle them. AI_COMPLETE takes images as well as text, and TO_FILE creates the stage-file reference that AI_COMPLETE and other file-accepting functions need. AI_TRANSCRIBE converts staged audio and video into text with timestamps and speaker information, and from there the text-analysis functions take over. AI_PARSE_DOCUMENT with extract_images gives you a document's embedded images, which you can then tag or analyze with AI_EXTRACT or AI_COMPLETE. AI_PARSE_DOCUMENT names multimodal RAG as one of its use cases.
For comparing content, AI_EMBED produces an embedding vector for text or an image. AI_MULTI_EMBED creates embeddings from text, images, audio or video, with segment metadata, for semantic search across modalities. AI_SIMILARITY calculates the embedding similarity between two inputs, which is useful for questions like "how close is this ticket to that known issue". The sources here give only that one-line description of AI_SIMILARITY. They do not list its supported input types or options.
Checkpoint 7 of 7· Match them up
Match each task to the function built for it
Tap a term, then the definition that fits it.
Each function's catalog entry names its job: AI_TRANSCRIBE for audio and video, AI_MULTI_EMBED for cross-modal embeddings, AI_SIMILARITY for pairwise similarity, and AI_COMPLETE for text-or-image generation.
“AI_TRANSCRIBE: Transcribes audio and video files stored in a stage, extracting text, timestamps, and speaker information.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.OCR mode is the right choice whenever documents contain tables, because it is the higher-quality mode.Why is that wrong?
OCR mode is for fast text extraction from scanned or text-heavy documents. Tables, headers, layout relationships and image extraction all need LAYOUT mode.
Covered in Turning files into text: AI_PARSE_DOCUMENT
2.Setting chunk_size to 512 in SPLIT_TEXT_RECURSIVE_CHARACTER gives you the recommended 512-token chunks.Why is that wrong?
chunk_size counts characters, not tokens. Since a token is about 4 characters, a 512-token chunk is roughly 2,000 characters.
Covered in Preparing text for retrieval: recursive and Markdown splitting
3.Picking the longest-context embedding model and indexing whole documents gives the best retrieval quality.Why is that wrong?
Smaller chunks usually give better retrieval and better downstream LLM answers, even when a long-context model is available.
Covered in Preparing text for retrieval: recursive and Markdown splitting
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“OCR mode is recommended for fast, high-quality text extraction from scanned or text-heavy documents such as contracts, insurance claims, and manuals.”
↩︎ Turning files into text: AI_PARSE_DOCUMENT“If using page_filter, page_split is implied, and you do not need to set it explicitly.”
↩︎ Turning files into text: AI_PARSE_DOCUMENT“AI_PARSE_DOCUMENT returns the content in Markdown format.”
↩︎ Turning files into text: AI_PARSE_DOCUMENT“Image understanding: Use extracted images with AI_EXTRACT or AI_COMPLETE for automatic tagging and analysis”
↩︎ Multimodal analytics: audio, images and similarity“LAYOUT mode is the preferred choice for most use cases, especially for complex documents.”
↩︎ Exam trap 1“It is optimized for extracting structured content such as tables, headers, and layout relationships, and is required for image extraction.”
↩︎ Checkpoint - 2.
“AI_EXTRACT uses arctic-extract, a proprietary vision-based large language model (LLM) that delivers high extraction accuracy.”
↩︎ Pulling structured fields out: AI_EXTRACT“content in a graphical form, such as logos, handwritten text (for example, signatures), tables, or checkmarks”
↩︎ Pulling structured fields out: AI_EXTRACT“Ask for a single value in each question.”
↩︎ Pulling structured fields out: AI_EXTRACT“The model for table extraction returns answers that are up to 4096 tokens long.”
↩︎ Pulling structured fields out: AI_EXTRACT“If the table covers several pages, split the document into multiple one-page documents, and join the results in postprocessing.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/sql-reference/functions/split_text_recursive_character-snowflake-cortexOfficial docs
“markdown: Separates on headers, code blocks, and tables, in addition to any separators in the separators field.”
↩︎ Preparing text for retrieval: recursive and Markdown splitting“Overlap is useful for ensuring that each chunk has some context about the previous chunk.”
↩︎ Preparing text for retrieval: recursive and Markdown splitting“An integer specifying the maximum number of characters in each chunk.”
↩︎ Exam trap 2 - 4.https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/cortex-search-overviewOfficial docs
“Cortex Search truncates the string to the size of the context window before embedding it into vector space for semantic search.”
↩︎ Preparing text for retrieval: recursive and Markdown splitting“Semantic reranking for reranking the most relevant documents in the result set.”
↩︎ Cortex Search: embedding models and semantic reranking“Retrieval augmented generation (RAG) is a technique for retrieving data from a knowledge base to enhance the generated response of a large language model.”
↩︎ Key concept“research shows that a smaller chunk size typically results in higher retrieval and downstream LLM response quality”
↩︎ Exam trap 3“Snowflake recommends splitting the text in your search column into chunks of no more than 512 tokens (about 385 English words).”
↩︎ Prediction“However, Cortex Search uses the full body of text for keyword-based retrieval.”
↩︎ Checkpoint - 5.https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/cortex-search-customize-scoringOfficial docs
“Disabling reranking reduces query latency by 100-300ms on average”
↩︎ Cortex Search: embedding models and semantic reranking“Reranking is not supported for batch search queries.”
↩︎ Cortex Search: embedding models and semantic reranking“While reranking can measurably increase result relevance, it can also noticeably increase query latency.”
↩︎ Checkpoint - 6.
“AI_COMPLETE: Generates a completion for a given text string or image using a selected LLM.”
↩︎ Multimodal analytics: audio, images and similarity“AI_SIMILARITY: Calculates the embedding similarity between two inputs.”
↩︎ Multimodal analytics: audio, images and similarity“AI_MULTI_EMBED: Creates multimodal embeddings from text, images, audio, or video, returning one or more vectors with segment metadata for semantic search across modalities.”
↩︎ Multimodal analytics: audio, images and similarity“AI_TRANSCRIBE: Transcribes audio and video files stored in a stage, extracting text, timestamps, and speaker information.”
↩︎ Checkpoint