What you will be able to do
- Explain how embedding-model context windows affect what Cortex Search indexes, and why Snowflake recommends 512-token chunks
- List the ways to query a Cortex Search Service, what multi-index queries change, and the privileges needed to create and query a service
- Describe a Cortex Knowledge Extension as a shared Cortex Search Service
- Describe how Cortex Analyst uses semantic views, their YAML specification, Semantic View Autopilot, verified queries and custom instructions
Key concept
Two kinds of grounding: retrieval and semantics — Snowflake's Gen AI features ground LLM answers in your data in two ways. Cortex Search retrieves relevant passages of unstructured text. Cortex Analyst generates SQL over structured data, guided by a semantic view. The other capabilities in this subdomain build on one or both of these.
1.Vector embeddings, tokens and context windows
Every Cortex model measures text in tokens. One token is about three quarters of an English word, or around four characters. The COUNT_TOKENS function counts the tokens in a string for a given model. This matters most for embedding models, which turn text into vectors so that semantically similar passages land close together. Each embedding model accepts a fixed number of tokens, called its context window. Cortex Search lets you pick a hosted model for its vector-search stage. These sources do not list the general Cortex AI function catalog, so this section covers embeddings as Cortex Search uses them.
| Model | Output dimensions | Context window (tokens) | Language support |
|---|---|---|---|
| snowflake-arctic-embed-m-v1.5 (default) | 768 | 512 | English-only |
| snowflake-arctic-embed-l-v2.0 | 1024 | 512 | Multilingual |
| snowflake-arctic-embed-l-v2.0-8k | 1024 | 8192 | Multilingual |
| voyage-multilingual-2 | 1024 | 32,000 | Multilingual |
If a value has more tokens than the context window allows, Cortex Search cuts it down to the window's size before embedding it. This happens both at indexing time and at serving time. Keyword retrieval is not truncated. A larger window is not automatically better, though. Snowflake recommends splitting text into chunks of no more than 512 tokens (about 385 English words), even though longer-context models exist. Smaller chunks make retrieval more precise and give a downstream LLM more relevant text. SPLIT_TEXT_RECURSIVE_CHARACTER is the built-in function for this chunking.
Checkpoint 1 of 6· Check yourself
A team indexes long documents with snowflake-arctic-embed-m-v1.5 and does not chunk them. Which statement is accurate?
Values longer than the context window are truncated before embedding, and keyword retrieval still uses the full text. The documentation describes no error and no automatic model switch.
“Cortex Search truncates the string to the size of the context window before embedding it into vector space for semantic search”Source: docs.snowflake.com
Sources1
2.Cortex Search: hybrid retrieval, ways to query, multi-index and access control
Cortex Search is a managed hybrid search engine over your text. Each query combines vector search for semantic similarity, keyword search for lexical matches, and a semantic reranking step. Snowflake handles embedding, infrastructure and index refreshes, which follow the same refresh behavior as dynamic tables and the service's TARGET_LAG. Its main uses are RAG for chatbots, enterprise search bars, and acting as the retrieval layer for Cortex Agents.
You can create a service with CREATE CORTEX SEARCH SERVICE or in Snowsight under AI & ML » Cortex Search. You can query it three ways: through the Python API, the REST API, or the SQL SEARCH_PREVIEW function. All three accept the same core parameters (query, columns, filter, limit, scoring options). Filters only work on columns listed in ATTRIBUTES.
SELECT PARSE_JSON(
SNOWFLAKE.CORTEX.SEARCH_PREVIEW(
'cortex_search_db.services.transcript_search_service',
'{
"query": "internet issues",
"columns":[
"transcript_text",
"region"
],
"filter": {"@eq": {"region": "North America"} },
"limit":1
}'
)
)['results'] as results;Multi-index queries. A service can index several columns. The SQL and Python APIs accept a multi_index_query parameter: a map from each indexed column name to text queries ({"text": ...}) or vector queries ({"vector": [...]}). Naming only the indexes you need gives more refined results and costs less, because fewer columns are searched. Multi-index services can still be searched through the REST API or without multi_index_query. In that case the search runs across every indexed column, which affects query cost.
Checkpoint 2 of 6· Check yourself
A developer queries a multi-index Cortex Search service through the REST API without multi_index_query. What happens?
REST and parameter-free queries still work, but they search every indexed column. Restricting the indexes with multi_index_query is how you cut cost.
“Multi-index Cortex Search services can still be searched through the REST API or without the multi_index_query parameter.”Source: docs.snowflake.com
Access control. The privileges for creating a service differ from those for querying one. Change tracking must be enabled on every underlying object. A service performs searches with owner's rights, using the same security model as other owner's-rights objects.
| Operation | What the role needs |
|---|---|
| Create a service | SNOWFLAKE.CORTEX_USER or SNOWFLAKE.CORTEX_EMBED_USER database role; CREATE CORTEX SEARCH SERVICE or OWNERSHIP on the schema; SELECT on underlying tables or views; USAGE on the refresh warehouse |
| Query a service | USAGE on the service, and on its database and schema |
| Suspend or resume with ALTER | OPERATE on the service |
Checkpoint 3 of 6· Fill the gap
Which privilege lets the customer_support role query the service?
GRANT ? ON CORTEX SEARCH SERVICE transcript_search_service TO ROLE customer_support;Querying needs USAGE on the service, plus USAGE on its database and schema. OPERATE is only for suspending and resuming the service.
Source: docs.snowflake.com3.Cortex Knowledge Extensions: sharing a search service
Cortex Knowledge Extensions (CKEs) are not a new kind of index. A CKE is a Cortex Search Service that a provider shares on the Snowflake Marketplace or through a private or organizational listing. The provider loads text into a table, creates a search service on it, and lists the service. A consumer then builds a Cortex AI application, such as a chatbot, with Cortex AI Functions or the Cortex Agent API. When a user sends a prompt, the application runs a semantic search against the CKE, and its LLM reasons over the results before answering with citations and attribution. For those citations to work, the provider should include a SOURCE_URL column in the indexed columns.
Providers get four features: content protection, management, trial support and monetization. Content protection limits the percentage of the indexed corpus a consumer can retrieve within a rolling 24-hour period. Once a consumer reaches that limit, further queries are blocked until it refreshes. Providers pay to host the service, including indexing, serving and replication. Consumers pay the provider if the CKE isn't free. CKEs are available in any region where Cortex Search is available.
Checkpoint 4 of 6· Check yourself
Which statement correctly describes a Cortex Knowledge Extension?
A CKE is a provider's Cortex Search Service shared through a listing. The provider hosts and pays for the index.
“Cortex Knowledge Extensions (CKEs) are Cortex Search Services that can be shared on the Snowflake Marketplace or via private listings or organizational listings.”Source: docs.snowflake.com
Sources3
4.Cortex Analyst and semantic views
Cortex Search covers unstructured text. Cortex Analyst covers structured data. It is a fully managed text-to-SQL service, exposed as a REST API, that answers business questions in natural language. It runs the SQL it generates in your virtual warehouse under your RBAC policies. A database schema alone does not capture business definitions, so Cortex Analyst relies on semantic views: schema-level objects that describe business concepts on top of your physical tables.
| Element | Purpose |
|---|---|
| Logical tables | Business entities such as customers, orders or products |
| Dimensions | Categorical context such as customer name or order date |
| Facts | Row-level quantitative data such as sale amounts |
| Metrics | Aggregations into KPIs such as total revenue |
| Relationships | How tables join together |
Semantic views also carry descriptions, synonyms, verified examples (sample questions paired with SQL that guide query generation), access modifiers that mark facts and metrics as public or private, and custom instructions that guide Cortex Analyst's SQL generation and how it categorizes questions.
A semantic view can be authored as a YAML specification. In Semantic Studio you define the view as a YAML file and deploy it to a live Snowflake object. You can also write it as code (DDL or YAML) for version control. Legacy semantic model YAML files stored on stages still work for backward compatibility, but semantic views are the recommended approach.
Semantic View Autopilot is the AI-assisted creation wizard in Snowsight (Workspaces » Add new » Semantic View). You describe what you want, and it generates logical tables, relationships and metrics. You can seed it with example SQL, a Tableau or Power BI file, or a YAML specification. It discards invalid example queries and adds the valid ones as verified queries. It also mines the creating role's query history for relationship and verified-query suggestions.
For access, use SNOWFLAKE.CORTEX_USER, which covers all Covered AI features, or SNOWFLAKE.CORTEX_ANALYST_USER, which covers only Cortex Analyst. CORTEX_USER is granted to PUBLIC by default.
Checkpoint 5 of 6· Check yourself
In Semantic View Autopilot you supply several example SQL queries as context. What happens to the valid ones?
Autopilot validates the example queries, extracts tables, columns and relationships from them, and keeps the valid queries as verified queries.
“Add valid queries to the semantic view as verified queries.”Source: docs.snowflake.com
Checkpoint 6 of 6· Exam question
Which practice best describes effective prompt design when calling AI_COMPLETE for a structured extraction task in Snowflake Cortex?
Correct answer: A — Provide a system-style instruction plus explicit output format constraints so the model returns structured, parseable text
- A. Pairing a clear instruction with explicit output constraints (such as a required JSON shape) is the standard way to get reliable, machine-parseable extraction results from a completion model.
- B. Passing raw values without any task framing leaves the model guessing at the desired structure, which increases the chance of inconsistent or unparseable output.
- C. Repeating the same text does not add task information and does not improve output structure; it only wastes prompt tokens.
- D. Removing formatting guidance increases output variability, which is the opposite of what a structured extraction task needs.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Text longer than the embedding model's context window is dropped from the search entirely.Why is that wrong?
Only the vector embedding is truncated to the context window. Keyword retrieval still uses the full text.
2.Consumers of a CKE pay to index the provider's content in their own account.Why is that wrong?
The provider hosts the Cortex Search Service and pays for indexing, serving and replication. Consumers pay the provider only if the CKE isn't free.
Covered in Cortex Knowledge Extensions: sharing a search service
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/cortex-search-overviewOfficial docs
“Snowflake recommends splitting the text in your search column into chunks of no more than 512 tokens (about 385 English words).”
↩︎ Vector embeddings, tokens and context windows“one token is equivalent to about 3/4 of an English word, or around 4 characters”
↩︎ Vector embeddings, tokens and context windows“the role of the querying user must have USAGE privileges on the service itself, as well as on the database and schema”
↩︎ Cortex Search: hybrid retrieval, ways to query, multi-index and access control“Change tracking must be enabled on all underlying objects used by a Cortex Search Service.”
↩︎ Cortex Search: hybrid retrieval, ways to query, multi-index and access control“However, Cortex Search uses the full body of text for keyword-based retrieval.”
↩︎ Exam trap 1“However, Cortex Search uses the full body of text for keyword-based retrieval.”
↩︎ Prediction“Cortex Search truncates the string to the size of the context window before embedding it into vector space for semantic search”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-search/query-cortex-search-serviceOfficial docs
“In addition, the SQL and Python APIs support multi-index queries.”
↩︎ Cortex Search: hybrid retrieval, ways to query, multi-index and access control“You can use three APIs for querying a Cortex Search Service”
↩︎ Cortex Search: hybrid retrieval, ways to query, multi-index and access control“Multi-index Cortex Search services can still be searched through the REST API or without the multi_index_query parameter.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-knowledge-extensions/cke-overviewOfficial docs
“Providers can limit the percentage of indexed content in their CKE that can be returned to their consumers within a rolling 24-hour period.”
↩︎ Cortex Knowledge Extensions: sharing a search service“Cortex Knowledge Extensions are available in any region where Cortex Search is available.”
↩︎ Cortex Knowledge Extensions: sharing a search service“Providers pay to host the Cortex Search Service in their account, including indexing, servicing, and replication to other regions.”
↩︎ Exam trap 2“Cortex Knowledge Extensions (CKEs) are Cortex Search Services that can be shared on the Snowflake Marketplace or via private listings or organizational listings.”
↩︎ Checkpoint - 4.
“Custom instructions: Provide guidance to Cortex Analyst for SQL generation and question categorization”
↩︎ Cortex Analyst and semantic views“Offering verified examples: Sample questions and their SQL answers guide query generation”
↩︎ Cortex Analyst and semantic views“Legacy semantic model YAML files (stored on stages) are still supported for backward compatibility, but Semantic Views are the recommended approach for new implementations.”
↩︎ Cortex Analyst and semantic views“CORTEX_USER provides access to all Covered AI features, while CORTEX_ANALYST_USER provides access only to Cortex Analyst.”
↩︎ Cortex Analyst and semantic views - 5.
“Instead of writing a YAML specification by hand, you describe what you want and Autopilot generates the logical tables, relationships, and metrics for you.”
↩︎ Cortex Analyst and semantic views“Add valid queries to the semantic view as verified queries.”
↩︎ Checkpoint - 6.
“You define a semantic view as a YAML file and deploy it to a live Snowflake object”
↩︎ Cortex Analyst and semantic views
Also cited
“They generate SQL over structured data using Cortex Analyst semantic views and use Cortex Search to retrieve insights from unstructured sources”
↩︎ Key concept