CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 12/56

    Evaluating Retrieval with AI Search Evaluation and Agent Evaluation Judges

    Use tools and metrics to evaluate retrieval performance

    8 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Describe the four-stage AI Search retrieval quality evaluation pipeline and its dashboard
    • Interpret common result patterns, such as hybrid vs ANN, reranker lift and wide confidence intervals
    • Pick the right Agent Evaluation / MLflow retrieval judge based on whether ground truth is available
    • Size and split an evaluation set the way Databricks recommends

    1.AI Search built-in retrieval evaluation

    Databricks says you need a reproducible evaluation system before you try to improve retrieval. The AI Search retrieval quality guide says that otherwise any tuning is guesswork. For AI Search, the fastest route is the built-in retrieval quality evaluation. It is in Beta and requires a managed Delta Sync AI Search index. You start it by clicking Evaluate search quality on the index page. You don't need to configure anything, because the defaults come from the index metadata. Any user with query access to the index can start a run, and that user owns the evaluation job.

    The evaluation runs a four-stage pipeline:

    1. Generate queries. It samples documents from the source table and has an LLM write a mix of natural-language and keyword queries. 2. Search across strategies. It runs each query with ANN, hybrid and full-text search, each with and without the reranker. 3. Score relevance. An LLM judge scores every query and document pair on the 0–3 scale. 4. Compute metrics and analyze. It computes the metrics with confidence intervals and saves the results so you can compare runs.

    Checkpoint 1 of 6· Put it in order

    Put the stages of the AI Search retrieval quality evaluation in order

    1. 1.Search each query across ANN, hybrid and full-text, with and without the reranker
    2. 2.Score each query and document pair with an LLM judge
    3. 3.Compute metrics with confidence intervals and persist the results
    4. 4.Generate queries from sampled source documents

    Sources12

    2.Reading the results dashboard

    Click View results to open the dashboard. Across the top are three summary indicators: the best DCG@10 score across all query types, the query type that achieved it (the recommended query type), and the number of queries evaluated. Below them is a bar chart of DCG@10 for each query type with and without the reranker. Next to it are tables of DCG@10 and average relevance. The dashboard also shows a line chart of average relevance by result position, the best- and worst-performing queries, and a comparison of base vs reranker results. It also lists failed queries, meaning queries whose top-1 result scored 0. A final chart tracks any chosen metric across runs over time.

    Every metric has a 95% confidence interval computed across per-query values. The interval shows whether a difference between strategies is real or just noise from a small query set. Wide intervals don't mean either strategy is bad. They mean you need more evaluation queries before comparing.

    Checkpoint 2 of 6· Check yourself

    On the AI Search dashboard, what makes a query appear in the failed-queries table?

    Sources1

    3.Turning results into decisions

    Running every strategy on the same query set makes the comparison direct. The docs map common result patterns to next steps. When the reranker helps, enabling it still depends on your latency budget.

    Common evaluation patterns and the suggested action
    PatternWhat it meansSuggested action
    Hybrid significantly better than ANNQueries benefit from keyword matchingUse hybrid search in production
    ANN approximately equal to hybridKeywords aren't adding value for your dataEither works; ANN is simpler
    Full-text significantly better than ANNEmbeddings may not capture your domain wellConsider fine-tuning the embedding model or using full-text search
    Reranker improves metrics significantlyCross-encoder provides meaningful quality liftEnable reranker if latency budget allows
    Wide confidence intervalsNot enough queries for reliable comparisonIncrease the number of evaluation queries
    All strategies score lowData quality or relevance issuesFollow the AI Search retrieval quality guide

    Checkpoint 3 of 6· Check yourself

    An evaluation shows full-text search scoring well above ANN for a pharmaceutical document index. What does the pattern suggest?

    Sources1

    4.Agent Evaluation and MLflow retrieval judges

    The AI Search tool evaluates the index on its own. To evaluate retrieval inside a full RAG chain, use Agent Evaluation, which is integrated with MLflow. It provides an evaluation harness and hosted LLM judges. Some retrieval metrics are deterministic: if your evaluation set lists the documents that contain the answer, document_recall can be computed directly. Others need a judge. For each judge, the question to ask is whether it needs ground truth (expectations).

    Retrieval-side judges and metrics, and what they require
    Judge / metricWhat it evaluatesNeeds ground truth?
    RetrievalRelevanceIs the retrieved context directly relevant to the user's request?No
    chunk_relevance/precisionWhat % of the retrieved chunks are relevant to the request?No
    RetrievalSufficiencyDoes the context provide all necessary information to produce the ground truth facts?Yes
    document_recallWhat % of the ground truth documents are represented in the retrieved chunks?Yes (deterministic)

    Response judges such as RetrievalGroundedness and Correctness evaluate the generated answer, not the retriever. They complement the judges above but don't replace them. Judges also need tuning: the cookbook says a judge must be tuned to the use case, by checking where it fails and improving it on those cases.

    Checkpoint 4 of 6· Check yourself

    You have queries but no labelled answers or documents yet. Which built-in judge can still tell you whether retrieved context is on-topic?

    Checkpoint 5 of 6· Exam question

    A technical support search tool must return the single correct troubleshooting article at the very top of the results, since agents only read the first couple of results before acting. Which metric best evaluates this requirement?

    Sources34

    5.Building the evaluation set

    Every metric above depends on the queries you run. The retrieval quality guide lists three ways to get them: an existing golden dataset of labelled query-answer pairs, a synthetic set generated from your documents, or ground-truth-free evaluation with Agent Evaluation judges. The data doesn't have to be perfect. What matters is comparing strategies against each other, not chasing absolute scores. A good set represents real production requests, includes hard cases, and is updated over time. It can be stored as a Delta table, and MLflow logs a snapshot of the version used in each evaluation.

    On size, Databricks recommends at least 30 questions and ideally 100–200. To avoid overfitting, split the set about 70% training (screening every experiment), 20% test (evaluating the best performers), and 10% validation (a final check before deploying to production).

    Checkpoint 6 of 6· Check yourself

    A team has a 150-question evaluation set and runs every retrieval experiment against all of it. What does Databricks recommend instead?

    Sources52

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If one strategy's DCG@10 is higher, it is the better strategy, even when the confidence intervals are wide.Why is that wrong?

      Wide confidence intervals mean there aren't enough queries to compare reliably. The fix is to add evaluation queries before choosing a strategy.

      Covered in Reading the results dashboard

    2. 2.If the reranker improves metrics, it should always be turned on.Why is that wrong?

      Reranking adds latency, so the suggested action is to enable it only if the latency budget allows.

      Covered in Turning results into decisions

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “No configuration is required, as default values are pre-populated based on your index metadata.”
      ↩︎ AI Search built-in retrieval evaluation
      “Any user with query access to the index can start an evaluation run and view the results dashboard.”
      ↩︎ AI Search built-in retrieval evaluation
      “All metrics include 95% confidence intervals computed across per-query values”
      ↩︎ Reading the results dashboard
      “Each strategy is also evaluated with and without the reranker.”
      ↩︎ Turning results into decisions
      “Increase the number of evaluation queries.”
      ↩︎ Exam trap 1
      “Enable reranker if latency budget allows.”
      ↩︎ Exam trap 2
      “The system samples documents from your source table and uses an LLM to generate realistic search queries.”
      ↩︎ Checkpoint
      “a table of failed queries (queries where the top-1 result was scored 0 (irrelevant))”
      ↩︎ Checkpoint
      “Consider fine-tuning your embedding model or using full-text search.”
      ↩︎ Checkpoint
    2. 2.
      “If you don't have evaluation in place, stop here and set it up first. Optimizing without measurement is guesswork.”
      ↩︎ AI Search built-in retrieval evaluation
      “Focus on relative improvements as you test different strategies, not absolute scores.”
      ↩︎ Building the evaluation set
    3. 3.
      “a subset of the retrieval metrics can also be computed deterministically.”
      ↩︎ Agent Evaluation and MLflow retrieval judges
      “For an LLM judge to be effective, it must be tuned to understand the use case.”
      ↩︎ Agent Evaluation and MLflow retrieval judges
    4. 4.
      “Does the context provide all necessary information to generate a response that includes the ground truth facts?”
      ↩︎ Agent Evaluation and MLflow retrieval judges
      “Is the retrieved context directly relevant to the user's request?”
      ↩︎ Checkpoint
    5. 5.
      “Databricks recommends at least 30 questions in your evaluation set, and ideally 100 - 200.”
      ↩︎ Building the evaluation set
      “Training set: ~70% of the questions.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.