What you will be able to do
- Map dataset columns and instrumented functions to the reserved attributes each evaluation metric needs
- Name the five built-in LLM-as-a-judge metrics and tell answer relevance apart from correctness
- Create, start, and score a run, and read its status correctly
- Compare runs across application versions in Snowsight and grant the privileges an evaluation needs
1.Datasets and reserved attributes
An evaluation tests an application version against a dataset. The dataset is a set of inputs, optionally with expected outputs (the ground truth). With the TruLens Python SDK you supply it as either a Snowflake table or a pandas dataframe. You map each dataset column to a reserved attribute. RECORD_ROOT.INPUT holds the prompt, RECORD_ROOT.INPUT_ID is an identifier (generated automatically if you leave it out), RETRIEVAL.QUERY_TEXT is the RAG query, and RECORD_ROOT.GROUND_TRUTH_OUTPUT is the expected answer. Your instrumented functions supply the output side: RETRIEVAL.RETRIEVED_CONTEXTS and RECORD_ROOT.OUTPUT.
@instrument(
span_type=SpanAttributes.SpanType.RETRIEVAL,
attributes={
SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
}
)
def retrieve_context(self, query: str) -> list:
return self.retrieve(query)Checkpoint 1 of 7· Check yourself
Which attribute comes from an instrumented function's return value rather than from a dataset column?
RETRIEVED_CONTEXTS is an output attribute captured from the retriever when the app runs. The other three are input attributes supplied by the dataset.
“Context retrieved from the search service or retriever.”Source: docs.snowflake.com
Sources1
2.Evaluation metrics: LLM-as-a-judge
The metrics use an LLM-as-a-judge approach. A Cortex LLM scores each output between 0 and 1 and explains its score. You can choose any LLM available in Cortex AI as the judge. If you don't choose one, llama3.1-70b is used. Under the hood the judge is called through AI_COMPLETE, so the role computing metrics must be allowed to call that function. A metric can only be scored when the attributes it needs were captured.
| Metric | What it judges | Required attributes |
|---|---|---|
| Context Relevance | Retrieved context is relevant to the query | RETRIEVAL.QUERY_TEXT, RETRIEVAL.RETRIEVED_CONTEXTS |
| Groundedness | Response is supported by the retrieved context (chain-of-thought) | RETRIEVAL.RETRIEVED_CONTEXTS, RECORD_ROOT.OUTPUT |
| Answer relevance | Response is relevant to the query; no ground truth | RECORD_ROOT.INPUT, RECORD_ROOT.OUTPUT |
| Correctness | Response aligns with the ground truth | RECORD_ROOT.INPUT, RECORD_ROOT.GROUND_TRUTH_OUTPUT, RECORD_ROOT.OUTPUT |
| Coherence | No logical gaps or contradictions | RECORD_ROOT.OUTPUT |
Two more measures come from the traces themselves, with no judge involved. Cost is calculated for each Cortex LLM call from the prompt_tokens and completion_tokens that AI_COMPLETE returns. Latency is measured for each instrumented function, rolled up into a total per input, and averaged across the whole run so you can compare configurations.
Checkpoint 2 of 7· Match them up
Match each metric to the one attribute set it needs
Tap a term, then the definition that fits it.
Correctness is the only metric that needs ground truth. Coherence looks only at the response, and the two retrieval metrics depend on the retrieved contexts.
“Correctness determines how aligned the generated response is with the ground truth.”Source: docs.snowflake.com
Sources1
3.Runs: creation, invocation, and status
A run is one batch evaluation of one application version against a dataset. You define it with RunConfig and add it with tru_app.add_run(). source_type is either DATAFRAME or TABLE. dataset_spec maps attributes to column names, label groups comparable runs, and llm_judge_name picks the judge model.
run_config = RunConfig(
run_name=run_name,
description="desc",
label="custom tag useful for grouping comparable runs",
source_type="DATAFRAME",
dataset_name="My test dataframe name",
dataset_spec={
"RETRIEVAL.QUERY_TEXT": "user_query_field",
"RECORD_ROOT.INPUT": "user_query_field",
"RECORD_ROOT.GROUND_TRUTH_OUTPUT": "golden_answer_field",
},
llm_judge_name="mistral-large2"
)
run = tru_app.add_run(run_config=run_config)Checkpoint 3 of 7· Put it in order
Put the four stages of a run in order
- 1.Computation: trigger metric computation by naming the metrics
- 2.Creation: add a run for an application version, specifying a dataset
- 3.Invocation: start the run so it calls the app for each input and stores traces
- 4.Visualization: review the results in Snowsight under AI & ML » Evaluations
A run has to produce outputs and traces before any metric can be computed. You view the results last.
“Computation: After invocation, trigger computation by specifying metrics to be computed.”Source: docs.snowflake.com
run.start() blocks until invocation and ingestion finish or time out. For a DATAFRAME source you pass the dataframe as input_df. The status then tells you how far the run has got. INVOCATION_COMPLETED means outputs and traces exist, but no scoring has happened yet. Scoring shows up as COMPUTATION_IN_PROGRESS and then COMPLETED. The PARTIALLY_ variants mean that some invocations or some metric computations failed.
Checkpoint 4 of 7· Check yourself
A run shows INVOCATION_COMPLETED, and Snowsight shows no groundedness scores. What does that mean?
INVOCATION_COMPLETED covers only the invocation stage. Metric scores appear after computation, when the status reaches COMPLETED.
“The run invocation completed with all outputs and traces created.”Source: docs.snowflake.com
Checkpoint 5 of 7· Exam question
What is the primary purpose of Snowflake AI Observability?
Correct answer: A — Trace executions and compute LLM-as-judge quality metrics for generative AI apps, built on TruLens
- A. AI Observability captures traces of an application's inputs, outputs, and intermediate steps and uses LLM-as-judge scoring to compute quality metrics, built on the open-source TruLens framework integrated with Snowflake.
- B. Row access policies are a data-governance feature unrelated to observability; they restrict row visibility for queries but do not trace or evaluate AI application behavior.
- C. Observability records traces and metrics for analysis by developers, but it does not automatically retrain or fine-tune the underlying LLMs based on feedback.
- D. Resource monitors control credit spend by suspending warehouses; observability is a separate capability focused on quality and tracing, not spend enforcement.
4.Computing server-side and custom metrics
After invocation reaches INVOCATION_COMPLETED or INVOCATION_PARTIALLY_COMPLETED, you can compute server-side metrics by name. compute_metrics() is asynchronous, and you can call it several times with different metric lists. You can't recompute a metric that has already been computed for the same run. Client-side metrics can sit in the same list. These are Metric objects that wrap any TruLens feedback function or any Python function, and you can tune them with your own criteria, examples, and score scale.
Checkpoint 6 of 7· Fill the gap
Which method scores this run on the built-in metrics?
run. ? (metrics=[
"coherence",
"answer_relevance",
"groundedness",
"context_relevance",
"correctness",
])run.compute_metrics() starts asynchronous LLM-judge scoring after invocation. start() only invokes the app.
Source: docs.snowflake.comSources2
5.Comparing runs and the privileges evaluations need
Comparison is the reason runs exist. You build one application version per mix of LLM, prompt, and parameters, then run each version against the same dataset. In Snowsight under AI & ML » Evaluations, open the External Agent, select several runs that share a dataset, and choose Compare. You see aggregated and record-level differences. Opening a single record shows its traces, span latency, and the judge's explanations. Deleting a run removes only its metadata. The records stay in the event table.
The required privileges follow the workflow. Using AI Observability needs the CORTEX_USER database role. Registering an app needs CREATE EXTERNAL AGENT on the schema. Creating and executing runs needs USAGE on the External Agent, CREATE TASK on the schema, and the global EXECUTE TASK privilege. Computing metrics needs USE AI FUNCTIONS, because the judge calls AI_COMPLETE. PUBLIC has this privilege by default, but it has to be granted explicitly if your account revoked it.
Checkpoint 7 of 7· Check yourself
A role can create runs, but compute_metrics fails after the account revoked a default grant from PUBLIC. Which privilege is most likely missing?
Metric computation calls AI_COMPLETE as the judge, so it needs USE AI FUNCTIONS. That grant normally comes through PUBLIC.
“AI Observability uses the AI_COMPLETE function as the LLM judge, so the role that runs the computation must be able to call AI_COMPLETE.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A high answer relevance score shows that the answer is correct.Why is that wrong?
Answer relevance doesn't use ground truth. Only correctness compares the response against an expected answer.
Covered in Evaluation metrics: LLM-as-a-judge
2.Deleting a run with run.delete() purges its traces from the event table.Why is that wrong?
Deleting a run removes only its metadata. The records it created stay stored in AI_OBSERVABILITY_EVENTS.
Covered in Comparing runs and the privileges evaluations need
3.If a metric result looks wrong, you can call compute_metrics again with the same metric to overwrite it.Why is that wrong?
You can call compute_metrics several times with different metrics, but a metric that has already been computed for a run can't be recomputed for that run.
Covered in Computing server-side and custom metrics
Practise it for real
Run a batch evaluation of a TruLens-instrumented RAG app and compare two versions in Snowsight
1.Have ACCOUNTADMIN grant your role SNOWFLAKE.CORTEX_USER, CREATE EXTERNAL AGENT and CREATE TASK on the app schema, EXECUTE TASK on the account, and USE AI FUNCTIONS on the account.
Why: Registering, running, and judging each need a different privilege.
You should see: Each GRANT statement succeeds.
2.Decorate the retrieval, generation, and entry-point methods with @instrument() span types, and map QUERY_TEXT and RETRIEVED_CONTEXTS on the retriever.
Why: Context relevance and groundedness need those attributes.
You should see: The app runs unchanged locally.
3.Register it with TruApp(app, app_name, app_version='v1', connector, main_method).
Why: This creates the External Agent and routes traces to AI_OBSERVABILITY_EVENTS.
You should see: The application appears under AI & ML » Evaluations.
4.Create a RunConfig with a dataset_spec, call tru_app.add_run(), then run.start(input_df=...).
Why: Invocation produces the outputs and traces that the judge needs.
You should see: run.get_status() reports INVOCATION_COMPLETED.
5.Call run.compute_metrics with groundedness, context_relevance, and correctness.
Why: LLM-judge scoring happens only after invocation.
You should see: Status moves through COMPUTATION_IN_PROGRESS to COMPLETED.
6.Repeat with app_version='v2' on the same dataset, then select both runs in Snowsight and choose Compare.
Why: Comparing runs on a shared dataset shows which version to deploy.
You should see: You see aggregated and record-level differences side by side.
Stuck? Get a nudge
If scores are missing, check that the attributes each metric needs were captured on the right spans.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“It can also contain a set of expected outputs (the ground truth).”
↩︎ Datasets and reserved attributes“Each column in the dataset must be mapped to one of the following reserved attributes”
↩︎ Datasets and reserved attributes“With this approach, an LLM is used to generate a score (between 0 and 1) with an explanation for the application’s output”
↩︎ Evaluation metrics: LLM-as-a-judge“If no LLM judge is specified, llama3.1-70b is used as the default judge.”
↩︎ Evaluation metrics: LLM-as-a-judge“Cost is calculated for each LLM invocation call that relies on Cortex LLMs based on the token usage information”
↩︎ Evaluation metrics: LLM-as-a-judge“The metric computation is completed with detailed outputs and traces.”
↩︎ Runs: creation, invocation, and status“You can compare the aggregated and record-level differences between the versions to identify improvements”
↩︎ Comparing runs and the privileges evaluations need“To register an application, your role must have CREATE EXTERNAL AGENT privileges on the schema.”
↩︎ Comparing runs and the privileges evaluations need“Note that this doesn’t rely on ground truth answer reference, and therefore this is not equivalent to assessing answer correctness.”
↩︎ Exam trap 1“Deleting a run deletes the metadata associated with the run. The records created as part of the run aren’t deleted and remain stored.”
↩︎ Exam trap 2“Context retrieved from the search service or retriever.”
↩︎ Checkpoint“Correctness determines how aligned the generated response is with the ground truth.”
↩︎ Checkpoint“Computation: After invocation, trigger computation by specifying metrics to be computed.”
↩︎ Checkpoint“The run invocation completed with all outputs and traces created.”
↩︎ Checkpoint“AI Observability uses the AI_COMPLETE function as the LLM judge, so the role that runs the computation must be able to call AI_COMPLETE.”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/user-guide/snowflake-cortex/ai-observability/evaluate-applications-trulensOfficial docs
“run.start() is blocking until invocation and ingestion complete or time out.”
↩︎ Runs: creation, invocation, and status“To compute server-side metrics after invocation status is INVOCATION_COMPLETED or INVOCATION_PARTIALLY_COMPLETED”
↩︎ Computing server-side and custom metrics“Client-side metrics can be any TruLens feedback function or any Python function.”
↩︎ Computing server-side and custom metrics“To compare runs that share a dataset, select multiple runs and choose Compare.”
↩︎ Comparing runs and the privileges evaluations need“a metric can’t be recomputed for the same run.”
↩︎ Exam trap 3