What you will be able to do
- Build a benchmark set from real user questions and explain how benchmark questions are run
- Choose between Chat mode and Agent mode for a benchmark run and interpret the Evaluations results
- Review an evaluation by comparing model output with ground truth, and handle questions that need manual review
- Keep Unity Catalog metadata and prompt-matching values current so Genie keeps generating correct SQL
1.Building a benchmark set
Fixing one response at a time tells you that one answer is now correct. It cannot tell you whether the Genie space is getting better overall, or whether yesterday's new instruction broke something else. Benchmarks answer that. A benchmark is a set of test questions you run to measure Genie's overall response accuracy, and each space can hold up to 500 of them. The most useful benchmark sets cover the questions users ask most often.
The questions should come from real usage. After you have reviewed a question in chat, the response's kebab menu has Add as benchmark. Databricks also suggests adding variations and different phrasings of the same question, so the benchmark tests whether Genie understands the intent and not just one wording.
Because each question runs as a new query, the only context Genie has is what is defined in the space: instructions, example SQL, and SQL functions. That is why benchmarks work as a regression test for your context changes. Run them after you update instructions, and any drop in accuracy points to the change that caused it.
Checkpoint 1 of 4· Check yourself
What does Genie use when it processes a benchmark question?
Each benchmark question is a new query that relies only on the context configured in the space. That is what makes it a test of that context.
“Each question is processed as a new query, using the instructions defined in the agent, including any provided example SQL and SQL functions.”Source: docs.databricks.com
2.Running benchmarks and reading the results
Benchmarks run in one of two modes, and you choose the mode when you run them, not when you add a question. Use the toggle at the top right of the Benchmarks tab.
| Mode | How accuracy is judged | Notes |
|---|---|---|
| Chat mode | Genie compares its SQL-generated results against a provided SQL answer | The default mode |
| Agent mode | An LLM judge grades the responses | Uses Agent mode's multi-step reasoning; an optional evaluation note can guide grading |
The Evaluations tab lists your runs. Each has an evaluation name (a timestamp), an execution status (completed, paused, or unsuccessful), an accuracy score across all benchmark questions, and the user who ran it. If a run includes questions with no predefined SQL answer, it is marked for review, and its accuracy score appears only after someone reviews those questions. In the detailed view, you compare each Model output with its Ground truth. Results rated Bad come with an explanation of how the output differs. You can change the assessment of any question. Response details stay in the evaluation for one week, so review a run soon after it finishes.
Checkpoint 2 of 4· Put it in order
Put the steps for reviewing an individual evaluation in order
- 1.Review and compare the Model output response with the Ground truth response
- 2.Click the timestamp in the Evaluation name column to open that run
- 3.Near the top of the Genie space, click Benchmark
- 4.Use the question list on the left to open a question
You open the Benchmarks area, choose a run by its timestamp, choose a question, and then compare the output with the ground truth.
“Review and compare the Model output response with the Ground truth response.”Source: docs.databricks.com
Checkpoint 3 of 4· Exam question
The finance director tells the Genie space owner that Genie's answers about "net revenue" consistently ignore returns and refunds, because analysts on the team have always defined that term to exclude them. This confusion shows up across many differently worded questions. What should the space owner update?
Correct answer: A — Update the space's general instructions to state the business definition of net revenue, subtracting returns and refunds.
- A. Because the misunderstanding spans many differently phrased questions, a general instruction that states the definition applies broadly, teaching Genie the business rule regardless of how a user asks.
- B. One example query only helps requests that closely match its exact phrasing, so it would leave the same misunderstanding in place for every other way users ask about net revenue.
- C. A benchmark measures whether responses are accurate against an expected SQL answer, but adding one without fixing instructions or assets does nothing to correct the underlying miscalculation.
- D. Renaming a production column is a schema change with downstream risk, and a name alone cannot convey the exclusion rule as reliably as an explicit written definition can.
Some benchmark questions in the run have no predefined SQL answer, so the run is marked for review. The accuracy measure appears only after those questions have been reviewed manually.
Sources1
3.Keeping Unity Catalog metadata current
Benchmarks and feedback often trace a wrong answer back to the data layer instead of the instructions. Genie writes SQL from Unity Catalog column names and descriptions, and the guidance calls quality table and column descriptions in Unity Catalog critical for accuracy. Improving and maintaining those descriptions at the source is therefore part of optimizing a space. If AI-generated descriptions are available, check them and use them only if they match what you would have written yourself.
There are two layers to keep in mind. In Configure > Sources, a table's default description is its Unity Catalog metadata. Editors can override it with a description specific to the space, and clicking Reset restores the Unity Catalog description. Edits made in the space, including synonyms and everything else in the knowledge store, stay inside the space and never change Unity Catalog. To fix the metadata for every consumer, change it in Unity Catalog itself.
Unity Catalog constraints matter too. Knowledge mining reads the Unity Catalog metadata of the tables in your space and automatically saves primary and foreign keys as join relationships. Without foreign keys in Unity Catalog, Genie may not know how to join tables. The fix is to define them in Unity Catalog where you can, or else to define join relationships in the knowledge store.
Data values change too. When new values appear in a column, or existing values change format, refresh that column's stored values from its kebab menu with Refresh prompt matching. If you skip this, Genie may filter on the wrong value, for example 'California' when the table stores 'CA'. The sources describe these two refresh actions (Reset to the Unity Catalog description, and Refresh prompt matching). They do not document a separate one-click sync of all Unity Catalog metadata.
| What changed | Action |
|---|---|
| New values were added to a column, or value formats changed | Kebab menu in the column view > Refresh prompt matching |
| A description overridden in the space should match Unity Catalog again | Click Reset to restore the Unity Catalog description |
| Primary and foreign keys are now defined in Unity Catalog | Knowledge mining saves them as join relationships in the space |
| The description itself is wrong for every consumer | Correct the table or column description in Unity Catalog; edits made in the space do not change it |
Checkpoint 4 of 4· Check yourself
A new product category was added to a table last week, and Genie now misses it when users filter by category. What should the space editor do first?
Prompt-matching data stores a column's values. Refreshing it after new values arrive lets Genie match user terms to the new category.
“Refreshing prompt matching data updates a column's stored values.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Each benchmark question is set to Chat mode or Agent mode when you add it.Why is that wrong?
The mode applies to a whole benchmark run and is chosen at run time with the toggle on the Benchmarks tab.
Covered in Running benchmarks and reading the results
2.Editing a table or column description in the Genie space updates that description in Unity Catalog.Why is that wrong?
Descriptions and synonyms set in the space apply only to that space. To change metadata for everyone, edit it in Unity Catalog.
Covered in Keeping Unity Catalog metadata current
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Each Genie Agent can contain up to 500 benchmark questions.”
↩︎ Building a benchmark set“A well-designed set of benchmarks covering the most frequently asked user questions helps evaluate the accuracy of your Genie Agent as you refine it.”
↩︎ Building a benchmark set“Chat mode: The default mode. Genie assesses accuracy by comparing its SQL-generated results against a provided SQL answer.”
↩︎ Running benchmarks and reading the results“For evaluation runs that require manual review, an accuracy measure appears only after those questions have been reviewed.”
↩︎ Running benchmarks and reading the results“The results of these responses appear in the evaluation details for one week.”
↩︎ Running benchmarks and reading the results“You select the mode when you run benchmarks, not when you add a question.”
↩︎ Exam trap 1“Benchmark questions run as new conversations. They do not carry the same context as a threaded Genie conversation.”
↩︎ Prediction“Each question is processed as a new query, using the instructions defined in the agent, including any provided example SQL and SQL functions.”
↩︎ Checkpoint“Review and compare the Model output response with the Ground truth response.”
↩︎ Checkpoint - 2.
“You can use variations and different question phrasings to test Genie's responses.”
↩︎ Building a benchmark set“Genie uses Unity Catalog column names and descriptions to generate responses.”
↩︎ Keeping Unity Catalog metadata current“This metadata is scoped to your Genie Agent and does not overwrite metadata stored in Unity Catalog.”
↩︎ Exam trap 2 - 3.
“The default table description shows the Unity Catalog metadata associated with your data asset.”
↩︎ Keeping Unity Catalog metadata current“Primary and foreign keys defined in your schema are automatically saved as join relationships in the Genie Agent.”
↩︎ Keeping Unity Catalog metadata current“Refreshing prompt matching data updates a column's stored values.”
↩︎ Checkpoint - 4.
“If foreign key references are not defined in Unity Catalog, your agent might not know how to join different tables together.”
↩︎ Keeping Unity Catalog metadata current“If new data has been added to relevant tables, refresh the values.”
↩︎ Keeping Unity Catalog metadata current