What you will be able to do
- Pick the Cortex AI Function that fits a categorization or text analytics step in a pipeline
- Build a SQL workflow that turns staged video, audio or documents into FILE objects and extracts structured results
- Explain how embeddings support semantic search and what a semantic view defines
- Estimate and monitor Cortex AI Function cost from token billing rules and Account Usage views
1.Cortex AI Functions: categorization and text analytics in SQL
Cortex AI Functions run LLM-powered analysis on text and images directly from SQL, and they are also available in Python. The models run inside the Snowflake service perimeter. To call them, a role needs the USE AI FUNCTIONS account-level privilege plus one of the CORTEX_USER or AI_FUNCTIONS_USER database roles.
Each function is built for a specific job, so you can drop it into a SELECT and use it as a step in a pipeline. For automated categorization, AI_CLASSIFY assigns text or images to categories you define, and AI_FILTER returns True or False so you can filter rows in WHERE or JOIN clauses. For text analytics, AI_SENTIMENT extracts sentiment, AI_SUMMARIZE condenses content, AI_TRANSLATE translates, AI_REDACT removes PII, and AI_EXTRACT pulls specific information out of text or files. AI_AGG and AI_SUMMARIZE_AGG work across many rows (for example, every support ticket in a table) and are not limited by the context window. AI_COMPLETE is the general-purpose function for most generative tasks.
| Pipeline task | Function |
|---|---|
| Label each record with a user-defined category | AI_CLASSIFY |
| Keep only rows that meet a natural-language condition | AI_FILTER |
| Score customer feedback | AI_SENTIMENT |
| Insights across many rows, without context-window limits | AI_AGG / AI_SUMMARIZE_AGG |
| Strip personally identifiable information | AI_REDACT |
| Free-form generation or analysis | AI_COMPLETE |
These functions are tuned for throughput rather than latency, which suits batch pipelines that process large tables. For interactive, latency-sensitive applications, Snowflake points you to the REST APIs instead: Complete, Embed and Agents.
Checkpoint 1 of 5· Check yourself
A nightly task tags 5 million support tickets with one of six team-defined categories. Which approach fits the documented guidance?
AI_CLASSIFY assigns user-defined categories, and AI Functions are designed for high-throughput batch work over SQL tables. The REST API is meant for latency-sensitive interactive use.
“Cortex AI Functions are optimized for throughput.”Source: docs.snowflake.com
Sources1
2.Multimodal workflows: extracting data from documents, images, audio and video
The same functions also accept media files. They process files on internal or external stages and extract insights from text, visual and audio signals. The bridge from a file to a function is the FILE data type. TO_FILE creates a reference to a staged file that AI_COMPLETE and other file-accepting functions can read, and a stage's directory table supplies the RELATIVE_PATH of every file. One CREATE TABLE AS statement therefore turns a whole folder of media into a table of FILE objects:
-- Create video table using TO_FILE data type
CREATE OR REPLACE TABLE video_ads_table AS
SELECT
TO_FILE('@video_ads', RELATIVE_PATH) AS video_file,
RELATIVE_PATH
FROM DIRECTORY(@video_ads);From that table, you choose a function according to the media type. When you call AI_COMPLETE with a response format, you can define the exact JSON schema to return, such as sentiment, detected brands, or a harmful-content flag, so the output arrives as structured data instead of free text. Note the release status: audio and video processing with AI_COMPLETE is in public preview, while the other multimodal capabilities are generally available.
| Media type | Primary functions | Common tasks |
|---|---|---|
| Documents | AI_COMPLETE, AI_PARSE_DOCUMENT, AI_EXTRACT | Q&A, summarization, extraction, chart understanding |
| Images | AI_COMPLETE, AI_CLASSIFY, AI_EMBED, AI_EXTRACT, AI_SIMILARITY, AI_FILTER | Caption, classify, extract entities, image search |
| Audio | AI_COMPLETE, AI_TRANSCRIBE | Transcribe, identify speakers, classify |
| Video | AI_COMPLETE, AI_TRANSCRIBE, AI_MULTI_EMBED (twelvelabs-marengo-embed-3-0 only) | Summarize, extract metadata, transcribe, semantic scene search |
AI_PARSE_DOCUMENT has two modes for documents: OCR mode extracts text, and LAYOUT mode extracts text with layout information. It can also pull out images embedded in a document. For speech, AI_TRANSCRIBE returns the text along with timestamps and speaker information. You can chain functions together: transcribe first, then pass the transcript to AI_COMPLETE, using PROMPT to build the prompt object.
Checkpoint 2 of 5· Fill the gap
Which function turns a staged podcast video into a text transcript?
SELECT ? (TO_FILE('@podcast_videos_S3', 'podcast-interview.mp4'));AI_TRANSCRIBE handles audio and video files on a stage and returns text with timestamps and speaker information. AI_PARSE_DOCUMENT works on documents.
Source: docs.snowflake.comSELECT
AI_COMPLETE('claude-sonnet-4-5',
PROMPT('Return a list of any Retail Brands mentioned in this podcast {0}',
TO_VARCHAR(transcription_results))) AS brands_identified
FROM podcast_video_transcription;Checkpoint 3 of 5· Exam question
A data platform team is building a custom document-management application that must store a stable, long-lived reference to each PDF's location alongside its metadata in a relational table, then resolve that reference to file bytes later through Snowflake's REST endpoint using a role that holds READ on the stage. Which URL type fits this design?
Correct answer: A — A permanent file URL produced by `BUILD_STAGE_FILE_URL`, since it identifies the database, schema, stage, and relative path and resolves through the REST endpoint whenever the caller's role holds the needed stage privilege.
- A. BUILD_STAGE_FILE_URL returns a permanent, non-expiring reference encoding the database, schema, stage, and path, and resolving it through the REST API still checks that the caller's role holds READ or USAGE on the stage, matching a long-lived stored reference with privilege-based access.
- B. A scoped URL is intentionally short-lived, tied to the generating user, and expires around the result cache period, so it is not suitable for a reference meant to be stored and resolved indefinitely later.
- C. A presigned URL deliberately skips Snowflake role checks entirely so it can be opened without authentication; it does not enforce or reflect stage-level role privileges at resolution time.
- D. RELATIVE_PATH is only a metadata string describing a file's location within the stage and is not accepted by the REST API as a substitute for one of the generated URL types.
3.Semantic analysis: embeddings, similarity and semantic views
Semantic analysis compares content by meaning, not exact words. AI_EMBED turns text or images into embedding vectors for similarity search, clustering and classification, and AI_SIMILARITY scores how similar two inputs are by their embeddings. For video, AI_MULTI_EMBED with the Twelve Labs Marengo 3 model converts each video into one or more 512-dimensional vectors that capture visual, audio and text meaning by segment. You store those vectors in a table and search them with SQL, for example to find every clip of someone riding a skateboard.
The exam guide also lists semantic views, which deal with meaning on the business side rather than in the files. A semantic view is a schema-level object that stores business concepts in the database: logical tables for entities such as customers or orders, relationships between them, and three kinds of definition. Facts are row-level attributes such as an individual sale amount. Metrics are aggregations such as Total Revenue. Dimensions are categorical attributes such as region or date. Defining "net revenue" once, with the correct aggregation, avoids dozens of inconsistent calculations spread across reports. You query a semantic view in a SELECT statement or use it from Cortex Agents, which read its definitions and generate SQL against the physical tables. You can create one with CREATE SEMANTIC VIEW, Semantic Studio in Workspaces, or a Snowsight wizard. Semantic views count as metadata.
Checkpoint 4 of 5· Match them up
Match each semantic view element to its role
Tap a term, then the definition that fits it.
Facts are row-level helpers, metrics aggregate them, dimensions give the viewing perspective, and logical tables model the business entities.
“Metrics are quantifiable measures of business performance calculated by aggregating facts or other columns from the same table”Source: docs.snowflake.com
4.Managing Cortex AI Function cost
Cortex AI Functions are billed by the number of tokens processed, at a per-function rate in credits per million tokens. As a rule of thumb, one token is about four characters of text. Warehouse cost applies on top of that, which is why the MEDIUM ceiling matters. The billing details vary by function, and the differences are easy to test:
| Function | Billing basis |
|---|---|
| AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_AGG, AI_SUMMARIZE, AI_TRANSLATE | Input and output tokens |
| AI_EMBED, AI_SIMILARITY | Input tokens only |
| AI_PARSE_DOCUMENT | Number of document pages processed |
| AI_EXTRACT | Input and output tokens; each document page counts as 970 input tokens |
| Audio input | 50 tokens per second of audio |
| AI_COUNT_TOKENS | Compute cost only, no token charge |
Some functions, including AI_CLASSIFY, AI_FILTER, AI_SENTIMENT and AI_TRANSLATE, add their own prompt to your input, so you are billed for more tokens than you supplied. AI_CLASSIFY also counts its labels, descriptions and examples as input tokens on every record, not once per call, so long category descriptions add up quickly on large tables. AI_COUNT_TOKENS lets you check prompt sizes before you spend.
To monitor spending, filter METERING_DAILY_HISTORY by service type, and use CORTEX_FUNCTIONS_USAGE_HISTORY for each function call or CORTEX_FUNCTIONS_QUERY_USAGE_HISTORY for each query. The per-query view has one row per model, so a query that calls two models shows two rows. Requests made through the REST API do not get this granular breakdown.
SELECT *
FROM SNOWFLAKE.ACCOUNT_USAGE.METERING_DAILY_HISTORY
WHERE SERVICE_TYPE='AI_SERVICES';Checkpoint 5 of 5· Check yourself
Which change most directly cuts AI_CLASSIFY token cost on a 10-million-row table?
Labels, descriptions and examples are billed as input tokens for every record processed, so trimming them reduces cost across all 10 million rows.
“AI_CLASSIFY labels, descriptions, and examples are counted as input tokens for each record processed, not just once for each AI_CLASSIFY call.”Source: docs.snowflake.com
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Scaling up the warehouse speeds up Cortex AI Functions, the same way it speeds up ordinary queries.Why is that wrong?
Snowflake recommends MEDIUM or smaller. A larger warehouse does not improve performance and only adds cost.
Covered in Managing Cortex AI Function cost
2.AI_PARSE_DOCUMENT is billed on input and output tokens, like AI_COMPLETE.Why is that wrong?
AI_PARSE_DOCUMENT is billed by the number of document pages processed.
Covered in Managing Cortex AI Function cost
3.Every multimodal capability of AI_COMPLETE is generally available.Why is that wrong?
Audio and video processing with AI_COMPLETE is in public preview. The other multimodal capabilities are GA.
Covered in Multimodal workflows: extracting data from documents, images, audio and video
Practise it for real
Transcribe a staged video with SQL and measure what the transcription cost
1.Create an internal stage with ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE') and DIRECTORY = ( ENABLE = true ), then upload a short .mp4 file with PUT.
Why: Server-side encryption keeps the files readable, and the directory table lists them for later steps.
You should see: The stage exists and the PUT reports the file as uploaded.
2.Refresh the directory metadata with ALTER STAGE ... REFRESH, then run SELECT * FROM DIRECTORY(@your_stage).
Why: The directory table only shows files after its metadata is refreshed.
You should see: One row with RELATIVE_PATH, SIZE, LAST_MODIFIED and FILE_URL for your video.
3.Create a table of FILE objects with TO_FILE('@your_stage', RELATIVE_PATH) selected FROM DIRECTORY(@your_stage).
Why: AI functions read staged media through FILE references.
You should see: A table with one FILE value per staged video.
4.Run AI_TRANSCRIBE on the FILE column.
Why: This extracts the speech in the video as text.
You should see: A JSON result that includes the audio_duration and text fields.
5.Query SNOWFLAKE.ACCOUNT_USAGE.CORTEX_FUNCTIONS_QUERY_USAGE_HISTORY WHERE query_id = '<your query id>'.
Why: This attributes credits and tokens to the specific query.
You should see: Token and credit figures for the transcription query, once the Account Usage data has arrived.
Stuck? Get a nudge
Audio is billed at 50 tokens per second, so a short clip keeps the test cheap.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“AI_CLASSIFY: Classifies text or images into user-defined categories.”
↩︎ Cortex AI Functions: categorization and text analytics in SQL“your role needs the USE AI FUNCTIONS account-level privilege and one of the CORTEX_USER or AI_FUNCTIONS_USER database roles”
↩︎ Cortex AI Functions: categorization and text analytics in SQL“For more interactive use cases where latency is important, use the REST API.”
↩︎ Cortex AI Functions: categorization and text analytics in SQL“AI_PARSE_DOCUMENT: Extracts text (using OCR mode) or text with layout information (using LAYOUT mode) from documents in an internal or external stage.”
↩︎ Multimodal workflows: extracting data from documents, images, audio and video“AI_EMBED: Generates an embedding vector for a text or image input, which can be used for similarity search, clustering, and classification tasks.”
↩︎ Semantic analysis: embeddings, similarity and semantic views“Cortex AI Functions are optimized for throughput.”
↩︎ Checkpoint - 2.
“These functions process files stored on internal or external stages, extracting insights from textual, visual, and audio signals.”
↩︎ Multimodal workflows: extracting data from documents, images, audio and video“You can define the exact schema you want the model to return”
↩︎ Multimodal workflows: extracting data from documents, images, audio and video“The Twelve Labs Marengo 3 embedding model converts each video into one or more 512-dimensional vectors”
↩︎ Semantic analysis: embeddings, similarity and semantic views“Audio and video processing using AI_COMPLETE is in public preview.”
↩︎ Exam trap 3 - 3.
“You can use Semantic Views in Cortex Agents and query these views in a SELECT statement.”
↩︎ Semantic analysis: embeddings, similarity and semantic views“You can store semantic business concepts directly in the database in a Semantic View, which is a schema-level object.”
↩︎ Semantic analysis: embeddings, similarity and semantic views“Metrics are quantifiable measures of business performance calculated by aggregating facts or other columns from the same table”
↩︎ Checkpoint - 4.
“Snowflake Cortex AI functions incur compute cost based on the number of tokens processed.”
↩︎ Managing Cortex AI Function cost“Audio files are billed at 50 tokens per second of audio.”
↩︎ Managing Cortex AI Function cost“You can’t get granular usage information for requests made with the REST API.”
↩︎ Managing Cortex AI Function cost“Snowflake recommends using a warehouse size no larger than MEDIUM when calling Snowflake Cortex AI Functions.”
↩︎ Exam trap 1“For AI_PARSE_DOCUMENT (or SNOWFLAKE.CORTEX.PARSE_DOCUMENT), billing is based on the number of document pages processed.”
↩︎ Exam trap 2“Using a larger warehouse than necessary does not increase performance, but can result in unnecessary costs.”
↩︎ Prediction“AI_CLASSIFY labels, descriptions, and examples are counted as input tokens for each record processed, not just once for each AI_CLASSIFY call.”
↩︎ Checkpoint