What you will be able to do
- Classify RAG retrieval and response metrics by whether they need ground truth, including the deterministic document_recall metric
- Explain why recall needs ground truth and precision does not
- Use the {{ expectations }} template variable to make a custom judge depend on ground truth, or leave it out to keep the judge reference-free
- Plan an evaluation that starts without labels and adds ground-truth judges later
1.RAG metrics: ground truth is not only about LLM judges
The Databricks RAG cookbook groups its recommended metrics into retrieval, response, cost and latency. It measures each one either deterministically or with an LLM judge. The ground-truth question cuts across both methods. The cookbook puts the LLM-judge split plainly: some judges, such as answer correctness, compare the human-labelled ground truth with the app outputs, and others, such as groundedness, do not need it.
| Dimension | Metric | Measured by | Needs ground truth? |
|---|---|---|---|
| Retrieval | chunk_relevance/precision | LLM judge | No |
| Retrieval | document_recall | Deterministic | Yes |
| Retrieval | context_sufficiency | LLM Judge | Yes |
| Response | correctness | LLM judge | Yes |
| Response | relevance_to_query | LLM judge | No |
| Response | groundedness | LLM judge | No |
| Response | safety | LLM judge | No |
| Cost | total_token_count, total_input_token_count, total_output_token_count | Deterministic | No |
| Latency | latency_seconds | Deterministic | No |
The retrieval pair shows the logic most clearly. Precision asks: of the chunks I retrieved, what share are relevant to the query? An LLM judge can decide relevance chunk by chunk, so you never need to know the full set of relevant documents. Recall asks: of all the documents known to be relevant, what share did I retrieve? That question has no answer unless someone has listed every relevant document beforehand. In the cookbook's example, two of the three retrieved results were relevant, so precision was 0.66. The retrieved results covered two of the four relevant documents, so recall was 0.5. In the MLflow dataset schema, this list is the expected_retrieved_context key, which the document_recall scorer reads.
The same pattern holds for the response metrics. Correctness and context sufficiency need ground truth; relevance, groundedness and safety do not. These match the Yes/No answers that the built-in judge table gives Correctness, RetrievalSufficiency, RelevanceToQuery, RetrievalGroundedness and Safety. Cost and latency come straight from the app's execution, so they never need labels.
Checkpoint 1 of 4· Check yourself
Your evaluation set has questions only, with no labelled documents. Which retrieval metric can you still compute?
Precision only asks whether each retrieved chunk is relevant to the request, which an LLM judge can decide. Recall and sufficiency both need ground truth.
“Computing precision does not require knowing all relevant items.”Source: docs.databricks.com
2.Custom judges: you decide whether ground truth is required
When the built-in judges don't fit, make_judge() lets you write the evaluation criteria in natural language. The instructions can reference only four template variables: {{ inputs }}, {{ outputs }}, {{ expectations }} and {{ trace }}. Custom names such as {{ question }} cause validation errors. Of the four, {{ expectations }} is the one documented as "Ground truths or expected outcomes". So whether a custom judge requires ground truth is not a fixed property of the judge type. It depends on whether your instructions reference {{ expectations }}.
from mlflow.genai.judges import make_judge
from typing import Literal
# Field-based judge: compare the response to the expected answer
correctness_judge = make_judge(
name="answer_correctness",
instructions=(
"Given the request in {{ inputs }}, decide whether the response in "
"{{ outputs }} matches the expected answer in {{ expectations }}.\n\n"
"Return 'correct' only if the response conveys the same facts as the "
"expected answer. Otherwise, return 'incorrect'."
),
feedback_value_type=Literal["correct", "incorrect"],
)The docs give the reference-free option directly: leave out {{ expectations }} when you don't have ground truth, for example to score tone or format from the inputs and outputs alone. The instructions must include at least one template variable, but they don't have to use all four. Trace-based judges work the same way. The documented tool-usage validator inspects only {{ trace }} to check that the agent picked appropriate tools with correct parameters, so it needs no labels. Note the contrast with the built-in ToolCallCorrectness judge, which compares against expectations.
Checkpoint 2 of 4· Check yourself
You need a custom judge that scores the tone of support replies, and you have no labelled answers. Which template should the instructions use?
Leaving out {{ expectations }} makes the judge reference-free. {{ question }} is not an allowed variable, and you only need to use at least one of the four.
“Omit {{ expectations }} when you don't have ground truth, for example to score tone or format from the inputs and outputs alone.”Source: docs.databricks.com
Checkpoint 3 of 4· Exam question
An agent invokes external tools such as a calculator and a search API. The evaluation team has a dataset specifying, for each test case, the exact tool names and arguments the agent should have called. They want a judge that flags cases where the agent's actual tool invocations differ from this expected sequence. Which judge requires this kind of ground truth?
Correct answer: A — ToolCallCorrectness, because it compares the agent's actual tool calls and arguments against the expected tool calls recorded for each case
- A. ToolCallCorrectness requires ground truth: it compares the agent's actual tool names and arguments against the expected tool calls recorded in the dataset for each test case, matching exactly what the team wants to check.
- B. ToolCallEfficiency judges whether the agent reached its result with an appropriate number of tool calls, evaluated independently of any recorded expected call sequence, so it does not rely on ground truth.
- C. Safety focuses on flagging potentially harmful outputs or actions and does not compare tool invocations to a recorded expected sequence, so it is reference-free rather than ground-truth based.
- D. RelevanceToQuery assesses whether an action addresses the user's request on its own terms, without checking it against a recorded expected tool-call sequence, so it does not require ground truth.
Sources3
3.Starting without labels, adding ground truth later
Knowing which judges need ground truth tells you what you can measure today. Databricks recommends a human-labelled evaluation set, but admits that curating labels takes time. Its advice is to start with an evaluation set that contains only questions and add ground-truth responses over time. Agent Evaluation can assess quality without ground truth, and it computes extra metrics such as answer correctness once ground truth exists.
In practice, your first runs use reference-free judges on unlabelled data. The harness example below scores pre-computed outputs that have no expectations at all:
# Evaluate pre-computed outputs
evaluation = mlflow.genai.evaluate(
data=results_data,
scorers=[Safety(), RelevanceToQuery()]
)As experts add expected_facts, expected_response or expected_retrieved_context to rows, you can add Correctness, RetrievalSufficiency and document_recall to the scorer list. Those additions are the metrics that tell you whether the answer was actually right, not just relevant and safe. The cookbook warns that you need both kinds of metric. A RAG app can answer poorly despite retrieving the correct context, and it can answer well despite faulty retrieval.
Checkpoint 4 of 4· Check yourself
A team's evaluation set has 50 questions and no labelled answers. What does the Databricks guidance say?
Agent Evaluation can assess quality without ground truth. Labels unlock additional metrics such as answer correctness, so the team can start now and add labels over time.
“Agent Evaluation can assess your chain's quality without ground truth, although, if ground truth is available, it computes additional metrics such as answer correctness.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Deterministic metrics never need ground truth; only LLM judges do.Why is that wrong?
document_recall is deterministic, and it still needs the labelled list of relevant documents to measure what share were retrieved.
Covered in RAG metrics: ground truth is not only about LLM judges
2.Every custom make_judge() judge needs ground truth because it is a custom evaluation.Why is that wrong?
A custom judge needs ground truth only if its instructions reference {{ expectations }}. You must use at least one of the four variables, not all of them.
Covered in Custom judges: you decide whether ground truth is required
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“Some LLM judges, such as answer correctness, compare the human-labeled ground truth vs. the app outputs.”
↩︎ RAG metrics: ground truth is not only about LLM judges“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
↩︎ RAG metrics: ground truth is not only about LLM judges“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
↩︎ Starting without labels, adding ground truth later“Computing recall requires your ground-truth to contain all relevant items.”
↩︎ Exam trap 1“document_recall | What % of the ground truth documents are represented in the retrieved chunks? | Deterministic | Yes”
↩︎ Prediction“Computing precision does not require knowing all relevant items.”
↩︎ Checkpoint - 2.
“expected_retrieved_context | document_recall scorer | Documents that should be retrieved”
↩︎ RAG metrics: ground truth is not only about LLM judges - 3.
“{{ expectations }} - Ground truths or expected outcomes”
↩︎ Custom judges: you decide whether ground truth is required“Analyze the {{ trace }} to verify correct tool usage.”
↩︎ Custom judges: you decide whether ground truth is required“Your instructions must include at least one template variable, but you don't need to use all of them.”
↩︎ Exam trap 2“Omit {{ expectations }} when you don't have ground truth, for example to score tone or format from the inputs and outputs alone.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-define-qualityOfficial docs
“You can get started by creating an evaluation set that only includes questions, and add the ground truth responses over time.”
↩︎ Starting without labels, adding ground truth later“Agent Evaluation can assess your chain's quality without ground truth, although, if ground truth is available, it computes additional metrics such as answer correctness.”
↩︎ Checkpoint