What you will be able to do
- Explain what ground truth means in an MLflow evaluation dataset and where it is stored
- Name the built-in single-turn judges that require ground truth: Correctness, RetrievalSufficiency and ToolCallCorrectness
- Tell apart look-alike judge pairs where one needs ground truth and the other does not
- Recognise that every built-in multi-turn conversation judge is reference-free
Key concept
Ground truth (expectations) — Ground truth is the human-supplied expected result for an evaluation row, such as the expected answer, the expected facts or the expected tool calls. In MLflow it is stored in the optional expectations field. A judge 'requires ground truth' when it can only score a response by comparing it against that field.
1.Where ground truth lives: the expectations field
Before you can say which judges need ground truth, you need to know what ground truth looks like to MLflow. An evaluation dataset row has a required inputs dictionary, which holds what you send to the app. It can also have an optional expectations dictionary, which holds what a human says the right result is. When you score pre-computed outputs or existing traces with mlflow.genai.evaluate(), expectations is still optional. The harness calls it "Optional ground truth for scorers."
Because the field is optional, you can run an evaluation without any labels at all. Judges that read only inputs and outputs will still produce scores. Judges that compare against an expected answer will not have the data they need. That difference is what the rest of this lesson is about.
Built-in judges look for specific keys inside expectations. The dataset reference lists reserved keys, and each one belongs to a particular consumer:
| Key | Used by | Holds |
|---|---|---|
| expected_facts | Correctness judge | List of facts that should appear |
| expected_response | Correctness judge | Exact or similar expected output |
| guidelines | Guidelines judge | Natural language rules to follow |
| expected_retrieved_context | document_recall scorer | Documents that should be retrieved |
Checkpoint 1 of 6· Check yourself
A dataset row has inputs and outputs but no expectations. Which statement is accurate?
inputs is required and expectations is optional. Reference-free judges still run on the row, but ground-truth judges such as Correctness have no expected answer to use.
“Optional ground truth for scorers”Source: docs.databricks.com
2.The built-in single-turn judges: which ones say "Yes"
Built-in LLM judges are predefined scorers that run on Databricks-hosted LLMs. Each one documents its arguments and whether it requires ground truth. The arguments column is the quickest tell. A judge that lists expectations among its arguments reads labelled data, and a judge that lists only inputs, outputs does not.
| Judge | Arguments | Requires ground truth |
|---|---|---|
| RelevanceToQuery | inputs, outputs | No |
| RetrievalRelevance | inputs, outputs | No |
| Safety | inputs, outputs | No |
| RetrievalGroundedness | inputs, outputs | No |
| Correctness | inputs, outputs, expectations | Yes |
| RetrievalSufficiency | inputs, outputs, expectations | Yes |
| Guidelines | inputs, outputs | No |
| ExpectationsGuidelines | inputs, outputs, expectations | No (but needs guidelines in expectations) |
| ToolCallCorrectness | inputs, outputs, expectations | Yes |
| ToolCallEfficiency | inputs, outputs | No |
Exactly three built-in single-turn judges answer "Yes". Correctness asks whether the response is correct compared with the provided ground truth. RetrievalSufficiency asks whether the retrieved context could support a response that includes the ground-truth facts. ToolCallCorrectness asks whether the tool calls and their arguments were correct for the query. Each one judges against a known right answer: the right facts, the right context or the right tool calls. That is why each needs a human to supply that answer first.
Checkpoint 2 of 6· Match them up
Match each built-in judge to its ground-truth requirement
Tap a term, then the definition that fits it.
Correctness and ToolCallCorrectness take expectations as an argument and are marked 'Yes'. RelevanceToQuery and Safety take only inputs and outputs.
“ToolCallCorrectness | inputs, outputs, expectations | Yes”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
A support-quality team has a labeled dataset of chat transcripts, each paired with a known correct answer written by subject-matter experts. They want an MLflow built-in judge that scores whether the agent's response matches this known correct answer for each transaction. Which judge should they use, and why does it require ground truth?
Correct answer: A — Correctness, because it compares the response against a provided expected_response or expected_facts to score factual accuracy
- A. Correctness is a built-in judge that requires ground truth: it compares the agent's response against an expected_response or expected_facts entry supplied in the evaluation dataset to score factual accuracy. This makes it the right fit when SMEs have already written known correct answers for each transcript.
- B. Safety evaluates whether content is free of harmful, offensive, or toxic material, and it does this without comparing against any reference answer. It would not tell the team whether a response matches the SME-written correct answer.
- C. RelevanceToQuery scores how directly a response addresses the user's request on its own, without needing a labeled reference. It cannot confirm factual match to a known correct answer the way a ground-truth judge can.
- D. RetrievalGroundedness checks whether the generated response is supported by the retrieved context rather than by an external labeled answer, so it does not rely on ground truth and would not validate against the SME-written correct answers.
Sources3
3.Look-alike pairs and the ExpectationsGuidelines edge case
Most mistakes come from judges whose names sound alike but sit on different sides of the line.
RetrievalGroundedness vs RetrievalSufficiency. Groundedness asks whether the response is grounded in the context, or whether the agent is hallucinating. It only needs the response and the context the app retrieved, so it is reference-free. Sufficiency asks whether that context contained enough to produce the ground-truth facts, so it needs expectations.
ToolCallEfficiency vs ToolCallCorrectness. Efficiency asks whether the tool calls avoided redundancy, which can be judged from the calls alone. Correctness asks whether the calls and their arguments were right, which needs expected tool calls.
Guidelines vs ExpectationsGuidelines. Guidelines applies one global set of natural-language rules to every row and takes no expectations. ExpectationsGuidelines applies per-row guidelines that domain experts labelled in an evaluation dataset. Its table entry reads "No (but needs guidelines in expectations)". It does not compare against a correct answer, but it still reads the expectations field, where the per-row guidelines are stored. That limits where it can run: the Guidelines page marks Guidelines as working in both offline evaluation and production monitoring, and ExpectationsGuidelines as for offline evaluation only.
Checkpoint 4 of 6· Fill the gap
Per-row guidelines for ExpectationsGuidelines go under which key?
from mlflow.genai.scorers import ExpectationsGuidelines
import mlflow
# Dataset with per-row guidelines
data = [
{
"inputs": {"question": "What is the capital of France?"},
"outputs": "The capital of France is Paris.",
" ? ": {
"guidelines": ["The response must be factual and concise"]
}
},ExpectationsGuidelines reads its per-row rules from the expectations dictionary. That is why it is reference-free and still cannot score a row without expectations.
Source: docs.databricks.comCheckpoint 5 of 6· Exam question
A RAG engineering team has a benchmark set where each question has expert-written expected_facts. They want to confirm the retriever pulled documents that actually contain enough information to support producing those expected facts, not merely documents that are topically relevant. Which built-in judge fits this need and requires ground truth?
Correct answer: A — RetrievalSufficiency, because it compares retrieved context against the expected_facts to confirm the necessary information was retrieved
- A. RetrievalSufficiency requires ground truth: it checks whether the retrieved context contains the information needed to produce the expected_facts recorded for that test case, which is exactly the completeness check the team wants.
- B. RetrievalRelevance scores how closely retrieved chunks relate to the query itself and does not compare against expected_facts, so it is reference-free and cannot confirm the context is sufficient to support a known set of facts.
- C. RetrievalGroundedness assesses whether the generated response is supported by the retrieved chunks, not whether the chunks contain enough information to match labeled facts, so it does not use ground truth in the way the team needs.
- D. Guidelines validates output against natural-language criteria that apply generally and does not compare against a per-case set of expected facts, so it would not confirm sufficiency of retrieved information relative to labeled facts.
4.Multi-turn conversation judges
For conversational agents, MLflow also provides judges that score a whole conversation instead of a single turn. They take a session argument; ConversationalGuidelines also takes guidelines. The judges are ConversationCompleteness, UserFrustration, KnowledgeRetention, ConversationalGuidelines, ConversationalRoleAdherence, ConversationalSafety and ConversationalToolCallEfficiency. All seven are marked "No" for ground truth. Each one assesses patterns visible in the conversation history itself, such as whether every question was addressed, whether the user got frustrated and whether the agent stayed in role. The docs say to use them both for evaluation during development and for monitoring in production.
Checkpoint 6 of 6· Check yourself
A team wants to check whether their assistant stays in its assigned role across a conversation, using unlabelled production traffic. What is accurate about ConversationalRoleAdherence?
Its only argument is session and it is marked 'No' for ground truth, so it runs without labels. The docs also say multi-turn judges can be used for monitoring in production.
“ConversationalRoleAdherence | session | No | Does the assistant maintain its assigned role throughout the conversation?”Source: docs.databricks.com
False. KnowledgeRetention takes only session and is marked 'No' for ground truth. It checks whether the agent correctly kept information from earlier in the same conversation, so the conversation itself is the reference.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.RetrievalGroundedness needs ground truth because it checks for hallucination against 'the truth'.Why is that wrong?
Groundedness checks the response against the context the app retrieved, not against a labelled answer. It takes only inputs and outputs and is marked 'No'.
Covered in Look-alike pairs and the ExpectationsGuidelines edge case
2.ExpectationsGuidelines is reference-free, so it can be dropped into production monitoring just like Guidelines.Why is that wrong?
It is marked 'No' for ground truth, but its per-row guidelines live in the expectations field, and the docs mark it for offline evaluation only.
Covered in Look-alike pairs and the ExpectationsGuidelines edge case
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“expectations has several reserved keys that are used by built-in LLM judges: guidelines, expected_facts, and expected_response.”
↩︎ Where ground truth lives: the expectations field“Ground truth labels, stored as a JSON-seralizable dict.”
↩︎ Key concept - 2.
“Optional ground truth for scorers”
↩︎ Where ground truth lives: the expectations field - 3.
“Is the response correct as compared to the provided ground truth?”
↩︎ The built-in single-turn judges: which ones say "Yes"“Correctness | inputs, outputs, expectations | Yes”
↩︎ The built-in single-turn judges: which ones say "Yes"“ExpectationsGuidelines | inputs, outputs, expectations | No (but needs guidelines in expectations)”
↩︎ Look-alike pairs and the ExpectationsGuidelines edge case“Use multi-turn judges both for evaluation during development and for monitoring in production.”
↩︎ Multi-turn conversation judges“KnowledgeRetention | session | No”
↩︎ Multi-turn conversation judges“RetrievalGroundedness | inputs, outputs | No”
↩︎ Exam trap 1“Does the context provide all necessary information to generate a response that includes the ground truth facts?”
↩︎ Prediction“ToolCallCorrectness | inputs, outputs, expectations | Yes”
↩︎ Checkpoint“ConversationalRoleAdherence | session | No | Does the assistant maintain its assigned role throughout the conversation?”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/mlflow3/genai/eval-monitor/concepts/judges/guidelinesOfficial docs
“Works in both offline evaluation and production monitoring.”
↩︎ Look-alike pairs and the ExpectationsGuidelines edge case“For offline evaluation only.”
↩︎ Exam trap 2