CertSafari
    Snowflake SnowPro Specialty: Gen AI (GES-C02)· Lessons

    Domain 3 · Lesson 11/15

    Evaluating AI Apps with TruLens: Metrics, Runs, and Comparisons

    Use Snowflake AI observability tools.

    10 min read
    7.25% of exam
    2 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Map dataset columns and instrumented functions to the reserved attributes each evaluation metric needs
    • Name the five built-in LLM-as-a-judge metrics and tell answer relevance apart from correctness
    • Create, start, and score a run, and read its status correctly
    • Compare runs across application versions in Snowsight and grant the privileges an evaluation needs

    1.Datasets and reserved attributes

    An evaluation tests an application version against a dataset. The dataset is a set of inputs, optionally with expected outputs (the ground truth). With the TruLens Python SDK you supply it as either a Snowflake table or a pandas dataframe. You map each dataset column to a reserved attribute. RECORD_ROOT.INPUT holds the prompt, RECORD_ROOT.INPUT_ID is an identifier (generated automatically if you leave it out), RETRIEVAL.QUERY_TEXT is the RAG query, and RECORD_ROOT.GROUND_TRUTH_OUTPUT is the expected answer. Your instrumented functions supply the output side: RETRIEVAL.RETRIEVED_CONTEXTS and RECORD_ROOT.OUTPUT.

    Mapping a retrieval function's parameter and return value to span attributes so context relevance can be scoredpython
    @instrument(
        span_type=SpanAttributes.SpanType.RETRIEVAL,
        attributes={
            SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
            SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
        }
    )
    def retrieve_context(self, query: str) -> list:
        return self.retrieve(query)

    Checkpoint 1 of 7· Check yourself

    Which attribute comes from an instrumented function's return value rather than from a dataset column?

    Sources1

    2.Evaluation metrics: LLM-as-a-judge

    The metrics use an LLM-as-a-judge approach. A Cortex LLM scores each output between 0 and 1 and explains its score. You can choose any LLM available in Cortex AI as the judge. If you don't choose one, llama3.1-70b is used. Under the hood the judge is called through AI_COMPLETE, so the role computing metrics must be allowed to call that function. A metric can only be scored when the attributes it needs were captured.

    Built-in metrics and the attributes each one requires
    MetricWhat it judgesRequired attributes
    Context RelevanceRetrieved context is relevant to the queryRETRIEVAL.QUERY_TEXT, RETRIEVAL.RETRIEVED_CONTEXTS
    GroundednessResponse is supported by the retrieved context (chain-of-thought)RETRIEVAL.RETRIEVED_CONTEXTS, RECORD_ROOT.OUTPUT
    Answer relevanceResponse is relevant to the query; no ground truthRECORD_ROOT.INPUT, RECORD_ROOT.OUTPUT
    CorrectnessResponse aligns with the ground truthRECORD_ROOT.INPUT, RECORD_ROOT.GROUND_TRUTH_OUTPUT, RECORD_ROOT.OUTPUT
    CoherenceNo logical gaps or contradictionsRECORD_ROOT.OUTPUT

    Two more measures come from the traces themselves, with no judge involved. Cost is calculated for each Cortex LLM call from the prompt_tokens and completion_tokens that AI_COMPLETE returns. Latency is measured for each instrumented function, rolled up into a total per input, and averaged across the whole run so you can compare configurations.

    Checkpoint 2 of 7· Match them up

    Match each metric to the one attribute set it needs

    Tap a term, then the definition that fits it.

    Sources1

    3.Runs: creation, invocation, and status

    A run is one batch evaluation of one application version against a dataset. You define it with RunConfig and add it with tru_app.add_run(). source_type is either DATAFRAME or TABLE. dataset_spec maps attributes to column names, label groups comparable runs, and llm_judge_name picks the judge model.

    Defining and adding a runpython
    run_config = RunConfig(
        run_name=run_name,
        description="desc",
        label="custom tag useful for grouping comparable runs",
        source_type="DATAFRAME",
        dataset_name="My test dataframe name",
        dataset_spec={
            "RETRIEVAL.QUERY_TEXT": "user_query_field",
            "RECORD_ROOT.INPUT": "user_query_field",
            "RECORD_ROOT.GROUND_TRUTH_OUTPUT": "golden_answer_field",
        },
        llm_judge_name="mistral-large2"
    )
    
    run = tru_app.add_run(run_config=run_config)

    Checkpoint 3 of 7· Put it in order

    Put the four stages of a run in order

    1. 1.Computation: trigger metric computation by naming the metrics
    2. 2.Creation: add a run for an application version, specifying a dataset
    3. 3.Invocation: start the run so it calls the app for each input and stores traces
    4. 4.Visualization: review the results in Snowsight under AI & ML » Evaluations

    run.start() blocks until invocation and ingestion finish or time out. For a DATAFRAME source you pass the dataframe as input_df. The status then tells you how far the run has got. INVOCATION_COMPLETED means outputs and traces exist, but no scoring has happened yet. Scoring shows up as COMPUTATION_IN_PROGRESS and then COMPLETED. The PARTIALLY_ variants mean that some invocations or some metric computations failed.

    Checkpoint 4 of 7· Check yourself

    A run shows INVOCATION_COMPLETED, and Snowsight shows no groundedness scores. What does that mean?

    Checkpoint 5 of 7· Exam question

    What is the primary purpose of Snowflake AI Observability?

    Sources21

    4.Computing server-side and custom metrics

    After invocation reaches INVOCATION_COMPLETED or INVOCATION_PARTIALLY_COMPLETED, you can compute server-side metrics by name. compute_metrics() is asynchronous, and you can call it several times with different metric lists. You can't recompute a metric that has already been computed for the same run. Client-side metrics can sit in the same list. These are Metric objects that wrap any TruLens feedback function or any Python function, and you can tune them with your own criteria, examples, and score scale.

    Checkpoint 6 of 7· Fill the gap

    Which method scores this run on the built-in metrics?

    run. ? (metrics=[
        "coherence",
        "answer_relevance",
        "groundedness",
        "context_relevance",
        "correctness",
    ])

    Sources2

    5.Comparing runs and the privileges evaluations need

    Comparison is the reason runs exist. You build one application version per mix of LLM, prompt, and parameters, then run each version against the same dataset. In Snowsight under AI & ML » Evaluations, open the External Agent, select several runs that share a dataset, and choose Compare. You see aggregated and record-level differences. Opening a single record shows its traces, span latency, and the judge's explanations. Deleting a run removes only its metadata. The records stay in the event table.

    The required privileges follow the workflow. Using AI Observability needs the CORTEX_USER database role. Registering an app needs CREATE EXTERNAL AGENT on the schema. Creating and executing runs needs USAGE on the External Agent, CREATE TASK on the schema, and the global EXECUTE TASK privilege. Computing metrics needs USE AI FUNCTIONS, because the judge calls AI_COMPLETE. PUBLIC has this privilege by default, but it has to be granted explicitly if your account revoked it.

    Checkpoint 7 of 7· Check yourself

    A role can create runs, but compute_metrics fails after the account revoked a default grant from PUBLIC. Which privilege is most likely missing?

    Sources21

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A high answer relevance score shows that the answer is correct.Why is that wrong?

      Answer relevance doesn't use ground truth. Only correctness compares the response against an expected answer.

      Covered in Evaluation metrics: LLM-as-a-judge

    2. 2.Deleting a run with run.delete() purges its traces from the event table.Why is that wrong?

      Deleting a run removes only its metadata. The records it created stay stored in AI_OBSERVABILITY_EVENTS.

      Covered in Comparing runs and the privileges evaluations need

    3. 3.If a metric result looks wrong, you can call compute_metrics again with the same metric to overwrite it.Why is that wrong?

      You can call compute_metrics several times with different metrics, but a metric that has already been computed for a run can't be recomputed for that run.

      Covered in Computing server-side and custom metrics

    Practise it for real

    Run a batch evaluation of a TruLens-instrumented RAG app and compare two versions in Snowsight

    1. 1.Have ACCOUNTADMIN grant your role SNOWFLAKE.CORTEX_USER, CREATE EXTERNAL AGENT and CREATE TASK on the app schema, EXECUTE TASK on the account, and USE AI FUNCTIONS on the account.

      Why: Registering, running, and judging each need a different privilege.

      You should see: Each GRANT statement succeeds.

    2. 2.Decorate the retrieval, generation, and entry-point methods with @instrument() span types, and map QUERY_TEXT and RETRIEVED_CONTEXTS on the retriever.

      Why: Context relevance and groundedness need those attributes.

      You should see: The app runs unchanged locally.

    3. 3.Register it with TruApp(app, app_name, app_version='v1', connector, main_method).

      Why: This creates the External Agent and routes traces to AI_OBSERVABILITY_EVENTS.

      You should see: The application appears under AI & ML » Evaluations.

    4. 4.Create a RunConfig with a dataset_spec, call tru_app.add_run(), then run.start(input_df=...).

      Why: Invocation produces the outputs and traces that the judge needs.

      You should see: run.get_status() reports INVOCATION_COMPLETED.

    5. 5.Call run.compute_metrics with groundedness, context_relevance, and correctness.

      Why: LLM-judge scoring happens only after invocation.

      You should see: Status moves through COMPUTATION_IN_PROGRESS to COMPLETED.

    6. 6.Repeat with app_version='v2' on the same dataset, then select both runs in Snowsight and choose Compare.

      Why: Comparing runs on a shared dataset shows which version to deploy.

      You should see: You see aggregated and record-level differences side by side.

    Stuck? Get a nudge

    If scores are missing, check that the attributes each metric needs were captured on the right spans.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “It can also contain a set of expected outputs (the ground truth).”
      ↩︎ Datasets and reserved attributes
      “Each column in the dataset must be mapped to one of the following reserved attributes”
      ↩︎ Datasets and reserved attributes
      “With this approach, an LLM is used to generate a score (between 0 and 1) with an explanation for the application’s output”
      ↩︎ Evaluation metrics: LLM-as-a-judge
      “If no LLM judge is specified, llama3.1-70b is used as the default judge.”
      ↩︎ Evaluation metrics: LLM-as-a-judge
      “Cost is calculated for each LLM invocation call that relies on Cortex LLMs based on the token usage information”
      ↩︎ Evaluation metrics: LLM-as-a-judge
      “The metric computation is completed with detailed outputs and traces.”
      ↩︎ Runs: creation, invocation, and status
      “You can compare the aggregated and record-level differences between the versions to identify improvements”
      ↩︎ Comparing runs and the privileges evaluations need
      “To register an application, your role must have CREATE EXTERNAL AGENT privileges on the schema.”
      ↩︎ Comparing runs and the privileges evaluations need
      “Note that this doesn’t rely on ground truth answer reference, and therefore this is not equivalent to assessing answer correctness.”
      ↩︎ Exam trap 1
      “Deleting a run deletes the metadata associated with the run. The records created as part of the run aren’t deleted and remain stored.”
      ↩︎ Exam trap 2
      “Context retrieved from the search service or retriever.”
      ↩︎ Checkpoint
      “Correctness determines how aligned the generated response is with the ground truth.”
      ↩︎ Checkpoint
      “Computation: After invocation, trigger computation by specifying metrics to be computed.”
      ↩︎ Checkpoint
      “The run invocation completed with all outputs and traces created.”
      ↩︎ Checkpoint
      “AI Observability uses the AI_COMPLETE function as the LLM judge, so the role that runs the computation must be able to call AI_COMPLETE.”
      ↩︎ Checkpoint
    2. 2.
      “run.start() is blocking until invocation and ingestion complete or time out.”
      ↩︎ Runs: creation, invocation, and status
      “To compute server-side metrics after invocation status is INVOCATION_COMPLETED or INVOCATION_PARTIALLY_COMPLETED”
      ↩︎ Computing server-side and custom metrics
      “Client-side metrics can be any TruLens feedback function or any Python function.”
      ↩︎ Computing server-side and custom metrics
      “To compare runs that share a dataset, select multiple runs and choose Compare.”
      ↩︎ Comparing runs and the privileges evaluations need
      “a metric can’t be recomputed for the same run.”
      ↩︎ Exam trap 3

    Ready to test yourself?

    Practise the 26 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.