CertSafari
    Snowflake SnowPro Advanced: Data Engineer (DEA-C02)· Lessons

    Domain 5 · Lesson 20/22

    Cortex AI Functions for Unstructured Data: Classification, Multimodal Extraction, Semantics and Cost

    Handle and process unstructured data.

    11 min read
    3.57% of exam
    4 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Pick the Cortex AI Function that fits a categorization or text analytics step in a pipeline
    • Build a SQL workflow that turns staged video, audio or documents into FILE objects and extracts structured results
    • Explain how embeddings support semantic search and what a semantic view defines
    • Estimate and monitor Cortex AI Function cost from token billing rules and Account Usage views

    1.Cortex AI Functions: categorization and text analytics in SQL

    Cortex AI Functions run LLM-powered analysis on text and images directly from SQL, and they are also available in Python. The models run inside the Snowflake service perimeter. To call them, a role needs the USE AI FUNCTIONS account-level privilege plus one of the CORTEX_USER or AI_FUNCTIONS_USER database roles.

    Each function is built for a specific job, so you can drop it into a SELECT and use it as a step in a pipeline. For automated categorization, AI_CLASSIFY assigns text or images to categories you define, and AI_FILTER returns True or False so you can filter rows in WHERE or JOIN clauses. For text analytics, AI_SENTIMENT extracts sentiment, AI_SUMMARIZE condenses content, AI_TRANSLATE translates, AI_REDACT removes PII, and AI_EXTRACT pulls specific information out of text or files. AI_AGG and AI_SUMMARIZE_AGG work across many rows (for example, every support ticket in a table) and are not limited by the context window. AI_COMPLETE is the general-purpose function for most generative tasks.

    Matching the pipeline task to a Cortex AI Function
    Pipeline taskFunction
    Label each record with a user-defined categoryAI_CLASSIFY
    Keep only rows that meet a natural-language conditionAI_FILTER
    Score customer feedbackAI_SENTIMENT
    Insights across many rows, without context-window limitsAI_AGG / AI_SUMMARIZE_AGG
    Strip personally identifiable informationAI_REDACT
    Free-form generation or analysisAI_COMPLETE

    These functions are tuned for throughput rather than latency, which suits batch pipelines that process large tables. For interactive, latency-sensitive applications, Snowflake points you to the REST APIs instead: Complete, Embed and Agents.

    Checkpoint 1 of 5· Check yourself

    A nightly task tags 5 million support tickets with one of six team-defined categories. Which approach fits the documented guidance?

    Sources1

    2.Multimodal workflows: extracting data from documents, images, audio and video

    The same functions also accept media files. They process files on internal or external stages and extract insights from text, visual and audio signals. The bridge from a file to a function is the FILE data type. TO_FILE creates a reference to a staged file that AI_COMPLETE and other file-accepting functions can read, and a stage's directory table supplies the RELATIVE_PATH of every file. One CREATE TABLE AS statement therefore turns a whole folder of media into a table of FILE objects:

    Registering staged videos as FILE objects using the stage's directory tablesql
    -- Create video table using TO_FILE data type
    CREATE OR REPLACE TABLE video_ads_table AS
    SELECT
        TO_FILE('@video_ads', RELATIVE_PATH) AS video_file,
        RELATIVE_PATH
    FROM DIRECTORY(@video_ads);

    From that table, you choose a function according to the media type. When you call AI_COMPLETE with a response format, you can define the exact JSON schema to return, such as sentiment, detected brands, or a harmful-content flag, so the output arrives as structured data instead of free text. Note the release status: audio and video processing with AI_COMPLETE is in public preview, while the other multimodal capabilities are generally available.

    Primary multimodal functions by media type
    Media typePrimary functionsCommon tasks
    DocumentsAI_COMPLETE, AI_PARSE_DOCUMENT, AI_EXTRACTQ&A, summarization, extraction, chart understanding
    ImagesAI_COMPLETE, AI_CLASSIFY, AI_EMBED, AI_EXTRACT, AI_SIMILARITY, AI_FILTERCaption, classify, extract entities, image search
    AudioAI_COMPLETE, AI_TRANSCRIBETranscribe, identify speakers, classify
    VideoAI_COMPLETE, AI_TRANSCRIBE, AI_MULTI_EMBED (twelvelabs-marengo-embed-3-0 only)Summarize, extract metadata, transcribe, semantic scene search

    AI_PARSE_DOCUMENT has two modes for documents: OCR mode extracts text, and LAYOUT mode extracts text with layout information. It can also pull out images embedded in a document. For speech, AI_TRANSCRIBE returns the text along with timestamps and speaker information. You can chain functions together: transcribe first, then pass the transcript to AI_COMPLETE, using PROMPT to build the prompt object.

    Checkpoint 2 of 5· Fill the gap

    Which function turns a staged podcast video into a text transcript?

    SELECT  ? (TO_FILE('@podcast_videos_S3', 'podcast-interview.mp4'));
    Chaining: running AI_COMPLETE over a stored transcriptsql
    SELECT
        AI_COMPLETE('claude-sonnet-4-5',
            PROMPT('Return a list of any Retail Brands mentioned in this podcast {0}',
                TO_VARCHAR(transcription_results))) AS brands_identified
    FROM podcast_video_transcription;

    Checkpoint 3 of 5· Exam question

    A data platform team is building a custom document-management application that must store a stable, long-lived reference to each PDF's location alongside its metadata in a relational table, then resolve that reference to file bytes later through Snowflake's REST endpoint using a role that holds READ on the stage. Which URL type fits this design?

    Sources21

    3.Semantic analysis: embeddings, similarity and semantic views

    Semantic analysis compares content by meaning, not exact words. AI_EMBED turns text or images into embedding vectors for similarity search, clustering and classification, and AI_SIMILARITY scores how similar two inputs are by their embeddings. For video, AI_MULTI_EMBED with the Twelve Labs Marengo 3 model converts each video into one or more 512-dimensional vectors that capture visual, audio and text meaning by segment. You store those vectors in a table and search them with SQL, for example to find every clip of someone riding a skateboard.

    The exam guide also lists semantic views, which deal with meaning on the business side rather than in the files. A semantic view is a schema-level object that stores business concepts in the database: logical tables for entities such as customers or orders, relationships between them, and three kinds of definition. Facts are row-level attributes such as an individual sale amount. Metrics are aggregations such as Total Revenue. Dimensions are categorical attributes such as region or date. Defining "net revenue" once, with the correct aggregation, avoids dozens of inconsistent calculations spread across reports. You query a semantic view in a SELECT statement or use it from Cortex Agents, which read its definitions and generate SQL against the physical tables. You can create one with CREATE SEMANTIC VIEW, Semantic Studio in Workspaces, or a Snowsight wizard. Semantic views count as metadata.

    Checkpoint 4 of 5· Match them up

    Match each semantic view element to its role

    Tap a term, then the definition that fits it.

    Sources123

    4.Managing Cortex AI Function cost

    Cortex AI Functions are billed by the number of tokens processed, at a per-function rate in credits per million tokens. As a rule of thumb, one token is about four characters of text. Warehouse cost applies on top of that, which is why the MEDIUM ceiling matters. The billing details vary by function, and the differences are easy to test:

    What gets billed, by function
    FunctionBilling basis
    AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_AGG, AI_SUMMARIZE, AI_TRANSLATEInput and output tokens
    AI_EMBED, AI_SIMILARITYInput tokens only
    AI_PARSE_DOCUMENTNumber of document pages processed
    AI_EXTRACTInput and output tokens; each document page counts as 970 input tokens
    Audio input50 tokens per second of audio
    AI_COUNT_TOKENSCompute cost only, no token charge

    Some functions, including AI_CLASSIFY, AI_FILTER, AI_SENTIMENT and AI_TRANSLATE, add their own prompt to your input, so you are billed for more tokens than you supplied. AI_CLASSIFY also counts its labels, descriptions and examples as input tokens on every record, not once per call, so long category descriptions add up quickly on large tables. AI_COUNT_TOKENS lets you check prompt sizes before you spend.

    To monitor spending, filter METERING_DAILY_HISTORY by service type, and use CORTEX_FUNCTIONS_USAGE_HISTORY for each function call or CORTEX_FUNCTIONS_QUERY_USAGE_HISTORY for each query. The per-query view has one row per model, so a query that calls two models shows two rows. Requests made through the REST API do not get this granular breakdown.

    Daily AI Services credit consumptionsql
    SELECT *
      FROM SNOWFLAKE.ACCOUNT_USAGE.METERING_DAILY_HISTORY
      WHERE SERVICE_TYPE='AI_SERVICES';

    Checkpoint 5 of 5· Check yourself

    Which change most directly cuts AI_CLASSIFY token cost on a 10-million-row table?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Scaling up the warehouse speeds up Cortex AI Functions, the same way it speeds up ordinary queries.Why is that wrong?

      Snowflake recommends MEDIUM or smaller. A larger warehouse does not improve performance and only adds cost.

      Covered in Managing Cortex AI Function cost

    2. 2.AI_PARSE_DOCUMENT is billed on input and output tokens, like AI_COMPLETE.Why is that wrong?

      AI_PARSE_DOCUMENT is billed by the number of document pages processed.

      Covered in Managing Cortex AI Function cost

    3. 3.Every multimodal capability of AI_COMPLETE is generally available.Why is that wrong?

      Audio and video processing with AI_COMPLETE is in public preview. The other multimodal capabilities are GA.

      Covered in Multimodal workflows: extracting data from documents, images, audio and video

    Practise it for real

    Transcribe a staged video with SQL and measure what the transcription cost

    1. 1.Create an internal stage with ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE') and DIRECTORY = ( ENABLE = true ), then upload a short .mp4 file with PUT.

      Why: Server-side encryption keeps the files readable, and the directory table lists them for later steps.

      You should see: The stage exists and the PUT reports the file as uploaded.

    2. 2.Refresh the directory metadata with ALTER STAGE ... REFRESH, then run SELECT * FROM DIRECTORY(@your_stage).

      Why: The directory table only shows files after its metadata is refreshed.

      You should see: One row with RELATIVE_PATH, SIZE, LAST_MODIFIED and FILE_URL for your video.

    3. 3.Create a table of FILE objects with TO_FILE('@your_stage', RELATIVE_PATH) selected FROM DIRECTORY(@your_stage).

      Why: AI functions read staged media through FILE references.

      You should see: A table with one FILE value per staged video.

    4. 4.Run AI_TRANSCRIBE on the FILE column.

      Why: This extracts the speech in the video as text.

      You should see: A JSON result that includes the audio_duration and text fields.

    5. 5.Query SNOWFLAKE.ACCOUNT_USAGE.CORTEX_FUNCTIONS_QUERY_USAGE_HISTORY WHERE query_id = '<your query id>'.

      Why: This attributes credits and tokens to the specific query.

      You should see: Token and credit figures for the transcription query, once the Account Usage data has arrived.

    Stuck? Get a nudge

    Audio is billed at 50 tokens per second, so a short clip keeps the test cheap.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “AI_CLASSIFY: Classifies text or images into user-defined categories.”
      ↩︎ Cortex AI Functions: categorization and text analytics in SQL
      “your role needs the USE AI FUNCTIONS account-level privilege and one of the CORTEX_USER or AI_FUNCTIONS_USER database roles”
      ↩︎ Cortex AI Functions: categorization and text analytics in SQL
      “For more interactive use cases where latency is important, use the REST API.”
      ↩︎ Cortex AI Functions: categorization and text analytics in SQL
      “AI_PARSE_DOCUMENT: Extracts text (using OCR mode) or text with layout information (using LAYOUT mode) from documents in an internal or external stage.”
      ↩︎ Multimodal workflows: extracting data from documents, images, audio and video
      “AI_EMBED: Generates an embedding vector for a text or image input, which can be used for similarity search, clustering, and classification tasks.”
      ↩︎ Semantic analysis: embeddings, similarity and semantic views
      “Cortex AI Functions are optimized for throughput.”
      ↩︎ Checkpoint
    2. 2.
      “These functions process files stored on internal or external stages, extracting insights from textual, visual, and audio signals.”
      ↩︎ Multimodal workflows: extracting data from documents, images, audio and video
      “You can define the exact schema you want the model to return”
      ↩︎ Multimodal workflows: extracting data from documents, images, audio and video
      “The Twelve Labs Marengo 3 embedding model converts each video into one or more 512-dimensional vectors”
      ↩︎ Semantic analysis: embeddings, similarity and semantic views
      “Audio and video processing using AI_COMPLETE is in public preview.”
      ↩︎ Exam trap 3
    3. 3.
      “You can use Semantic Views in Cortex Agents and query these views in a SELECT statement.”
      ↩︎ Semantic analysis: embeddings, similarity and semantic views
      “You can store semantic business concepts directly in the database in a Semantic View, which is a schema-level object.”
      ↩︎ Semantic analysis: embeddings, similarity and semantic views
      “Metrics are quantifiable measures of business performance calculated by aggregating facts or other columns from the same table”
      ↩︎ Checkpoint
    4. 4.
      “Snowflake Cortex AI functions incur compute cost based on the number of tokens processed.”
      ↩︎ Managing Cortex AI Function cost
      “Audio files are billed at 50 tokens per second of audio.”
      ↩︎ Managing Cortex AI Function cost
      “You can’t get granular usage information for requests made with the REST API.”
      ↩︎ Managing Cortex AI Function cost
      “Snowflake recommends using a warehouse size no larger than MEDIUM when calling Snowflake Cortex AI Functions.”
      ↩︎ Exam trap 1
      “For AI_PARSE_DOCUMENT (or SNOWFLAKE.CORTEX.PARSE_DOCUMENT), billing is based on the number of document pages processed.”
      ↩︎ Exam trap 2
      “Using a larger warehouse than necessary does not increase performance, but can result in unnecessary costs.”
      ↩︎ Prediction
      “AI_CLASSIFY labels, descriptions, and examples are counted as input tokens for each record processed, not just once for each AI_CLASSIFY call.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 13 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.