CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 3 · Lesson 10/15

    Cortex Cost Drivers: Tokens, Search Indexing and Serving, Agents and Analyst

    Manage, monitor, and optimize Snowflake Cortex costs.

    10 min read
    7.25% of exam
    4 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Say which unit each Cortex feature bills by: tokens, pages, messages or GB-month
    • Find the hidden token costs in AI function calls and use AI_COUNT_TOKENS to estimate a job before running it
    • Tell apart the cost categories of a Cortex Search Service and pick the setting that reduces each one
    • Cap how many tokens a Cortex Agent can use with the orchestration budget

    Key concept

    Metered unit per Cortex feature — Every Cortex feature bills credits against its own unit. AI functions bill by tokens (or pages for documents), Cortex Analyst bills by messages, and Cortex Search bills by tokens embedded plus GB-months of indexed data that it serves. To cut a cost, you first find the unit that drives it.

    1.AI functions: what counts as a billable token

    Cortex AI functions charge credits per million tokens. The rate for each function is listed in the Snowflake Service Consumption Table. A token is roughly four characters of text, but the count varies by model. The part that catches people out is which tokens count. Functions that generate text, such as AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_AGG, AI_SUMMARIZE and AI_TRANSLATE, bill both input and output tokens. Embedding functions (AI_EMBED, AI_SIMILARITY, SNOWFLAKE.CORTEX.EMBED_*) bill input tokens only. Cortex Guard bills input tokens on top of the AI_COMPLETE call, and its input is the output that AI_COMPLETE produced.

    Some tokens never appear in your query. AI_CLASSIFY, AI_FILTER, AI_AGG, AI_SENTIMENT, SUMMARIZE, TRANSLATE and several other functions add their own prompt to your text, so the billed count is higher than the text you sent. AI_CLASSIFY is a sharper case: its labels, descriptions and examples are billed again for every row. On a table with millions of rows, long label descriptions become a large share of the bill. Fewer tokens therefore means short outputs, short label sets, and only the columns the task needs.

    How each type of AI function call is billed
    FunctionWhat is billed
    AI_COMPLETE, AI_CLASSIFY, AI_FILTER, AI_AGG, AI_SUMMARIZE, AI_TRANSLATEInput and output tokens
    AI_EMBED, AI_SIMILARITY, SNOWFLAKE.CORTEX.EMBED_*Input tokens only
    AI_PARSE_DOCUMENTNumber of document pages processed
    AI_EXTRACTInput and output tokens; each document page counts as 970 tokens
    Audio input50 tokens per second of audio
    AI_COUNT_TOKENSCompute to run the function only, no token-based cost

    The last row of the table matters for planning. AI_COUNT_TOKENS charges no token-based cost, so you can run it over a sample to estimate a job's token volume before you pay for any inference. The warehouse also still costs money: the query that calls the function keeps a warehouse running. Snowflake recommends MEDIUM or smaller, because a bigger warehouse does not make AI functions run faster.

    Checkpoint 1 of 5· Match them up

    Match each function to what it is billed on

    Tap a term, then the definition that fits it.

    Checkpoint 2 of 5· Exam question

    A platform team runs AI_COMPLETE against a 500-row table every night to generate short product summaries. Monthly Cortex spend has grown far faster than the row count. Which factor is the most likely primary driver of the token-based charge for this workload?

    Sources1

    2.Cortex Search: six places credits go

    A Cortex Search Service is a pipeline, not one function call, and each stage bills separately. Indexing is billed when data changes. Serving is billed for as long as the service is up.

    Cortex Search cost categories and what drives each one
    CategoryWhat drives the cost
    Virtual warehouse computeRefresh queries against base objects, orchestrating embedding jobs, building the index
    EMBED_TEXT tokens computeTokens embedded for each inserted or updated row in the ON search column
    Serving computeGB/mo of uncompressed indexed data, charged while the service is available
    StorageFlat rate per TB for the materialized source query and serving structures
    Cloud services computeChange detection; billed only if above 10% of daily warehouse cost

    Embedding is incremental: you pay only for rows that were added or changed. The initial build is still a big one-off cost. Here is the documentation's worked example for 10 million rows of 500 tokens each:

    Embedding cost of the initial refreshtext
    (0.05 credits per 1 million tokens) * (10,000,000 rows) * (500 tokens per row) / (1,000,000 tokens)= 250 credits

    Serving is the cost that keeps coming back. It grows with the size of the indexed data and with the dimensions of the embedding model, because every vector is stored at 4 bytes per dimension. In multi-index services, each extra vector column multiplies that term, so larger vectors or more indexed columns cost more.

    Checkpoint 3 of 5· Check yourself

    Which Cortex Search cost is charged per GB per month of uncompressed indexed data?

    Sources2

    3.Reducing Cortex Search indexing and serving costs

    Each cost category has its own fix. For warehouse compute, keep the warehouse small: most of the build time goes to embedding, which does not speed up with more cores, so most services need MEDIUM and few gain anything above LARGE. Set target lag to how fresh the business actually needs the data, and suspend indexing during periods when freshness does not matter. Defining primary keys lets the service take an optimized refresh path when only a few rows have changed.

    For embedding tokens, the rule is to avoid re-embedding. Changing any value in a row re-embeds that row's search column, even if the search column itself did not change. So collect all changes to a row into one update. Each update also has a fixed cost, which makes fewer, larger updates cheaper. Any schema change to the source query triggers a full refresh. So does rebuilding the source table with CREATE OR REPLACE. Update the table incrementally with MERGE INTO instead, and add spare payload columns up front so you don't have to change the schema later.

    For serving, suspend the service during predictable idle periods. For services with irregular traffic, set AUTO_SUSPEND. To get a baseline for warehouse cost, first run the service on a dedicated warehouse so its refresh consumption can be isolated.

    Checkpoint 4 of 5· Check yourself

    A pipeline rebuilds the Cortex Search source table every night with CREATE OR REPLACE, although only about 1% of rows change. Which change most directly cuts embedding cost?

    Sources2

    4.Cortex Agents token budgets and Cortex Analyst message billing

    A Cortex Agent can be capped directly in its specification. The top-level orchestration object follows the OrchestrationConfig schema, and that schema includes a budget with two fields: seconds (a time budget) and tokens (a token budget). This limits a single agent run. Spend across all runs is governed by tag-based resource budgets, which belong to cost monitoring.

    Budget object in an agent's orchestration configurationjson
    {
      "seconds": 30,
      "tokens": 16000
    }

    Cortex Analyst bills by a different unit. It does not charge per token. It charges by the number of messages processed, at the rate in the Service Consumption Table. Its usage view records credits and REQUEST_COUNT per user in one-hour increments. To lower Analyst cost, reduce the number of messages, not the number of words in each one.

    Checkpoint 5 of 5· Check yourself

    Which quantity sets the credit charge for Cortex Analyst?

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.AI_CLASSIFY bills its label list once per call, so adding detailed label descriptions costs almost nothing.Why is that wrong?

      Labels, descriptions and examples are billed as input tokens for every row processed.

      Covered in AI functions: what counts as a billable token

    2. 2.An idle Cortex Search Service costs nothing because no queries are being served.Why is that wrong?

      Serving compute is charged while the service is available, even when it serves no queries. Suspend it or set AUTO_SUSPEND.

      Covered in Cortex Search: six places credits go

    3. 3.Only changes to the search column cause a row to be re-embedded.Why is that wrong?

      A change to any value in the row re-embeds that row's search column, so changes should be bundled into one update.

      Covered in Reducing Cortex Search indexing and serving costs

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “both input and output tokens are billable.”
      ↩︎ AI functions: what counts as a billable token
      “As a result, the billed token count is higher than the number of tokens in the text you provide.”
      ↩︎ AI functions: what counts as a billable token
      “AI_COUNT_TOKENS incurs only compute cost to run the function. No additional token-based costs are incurred.”
      ↩︎ AI functions: what counts as a billable token
      “Snowflake recommends using a warehouse size no larger than MEDIUM when calling Snowflake Cortex AI Functions.”
      ↩︎ AI functions: what counts as a billable token
      “Snowflake Cortex AI functions incur compute cost based on the number of tokens processed.”
      ↩︎ Key concept
      “AI_CLASSIFY labels, descriptions, and examples are counted as input tokens for each record processed, not just once for each AI_CLASSIFY call.”
      ↩︎ Exam trap 1
      “For AI_PARSE_DOCUMENT (or SNOWFLAKE.CORTEX.PARSE_DOCUMENT), billing is based on the number of document pages processed.”
      ↩︎ Checkpoint
    2. 2.
      “the embedding cost is only incurred for added or changed documents”
      ↩︎ Cortex Search: six places credits go
      “the size of the total indexed data, and the dimensionality of the selected vector embedding model”
      ↩︎ Cortex Search: six places credits go
      “Larger embedding vectors or higher numbers of index columns incur higher costs.”
      ↩︎ Cortex Search: six places credits go
      “Most services do not see improved indexing performance beyond a LARGE warehouse and many need only MEDIUM.”
      ↩︎ Reducing Cortex Search indexing and serving costs
      “fewer, bigger updates are less expensive than more frequent, smaller updates”
      ↩︎ Reducing Cortex Search indexing and serving costs
      “Defining primary keys on your Cortex Search Service can result in significant reductions to both the cost and latency of indexing.”
      ↩︎ Reducing Cortex Search indexing and serving costs
      “set the AUTO_SUSPEND property so that serving is suspended after a period of inactivity and resumed when the next query arrives”
      ↩︎ Reducing Cortex Search indexing and serving costs
      “You incur these costs while the service is available to respond to queries, even if no queries are served during a given period.”
      ↩︎ Exam trap 2
      “any change to any value within a row triggers the search column in that row to be embedded again”
      ↩︎ Exam trap 3
      “You incur these costs while the service is available to respond to queries, even if no queries are served during a given period.”
      ↩︎ Prediction
      “The compute cost for this component is incurred per GB per month (GB/mo) of uncompressed indexed data”
      ↩︎ Checkpoint
      “with a CREATE OR REPLACE command causes the service to fully refresh and embed all vectors again.”
      ↩︎ Checkpoint
    3. 4.
      “The information in the view includes the number of credits consumed each time Cortex Analyst is called, aggregated in one-hour increments.”
      ↩︎ Cortex Agents token budgets and Cortex Analyst message billing
      “Credit rate usage is based on the number of messages processed, as outlined in the Snowflake Service Consumption Table.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Monitoring Cortex Spend: Usage Views, Compute Pools, Alerts and Tag-Based Budgets

    Spotted a mistake, or was something unclear? Tell us.