What you will be able to do
- Separate retrieval problems from response problems when you review a RAG answer
- Match common response issues (irrelevance, hallucination, unsafe content, incorrectness, guideline violations) to the MLflow built-in judge that detects each one
- Tell which assessments need a ground-truth expectation and which do not
- Recognise quality and safety issues that only show up across a whole conversation
Key concept
Quality dimension (as assessed by an LLM judge) — A named part of response quality, such as relevance, groundedness, safety or correctness, that you assess on its own. Each one has a question you can ask about a response, and an LLM judge can answer it at scale. If you break 'is this answer good?' into separate dimensions, you can say exactly what went wrong.
1.Where response problems come from
When you assess a response qualitatively, the first job is to find where the issue started. Databricks splits RAG application performance into three dimensions. Retrieval quality asks whether the app fetched relevant supporting data. Response quality asks how well the app answered the user: whether the answer matches the ground truth, whether it is grounded in the retrieved context (or hallucinated), and whether it is safe. The source puts safety as 'how safe the response was (in other words, no toxicity)'. System performance covers cost and latency, such as token consumption and overall latency.
The first two are tied together, but not in the way people expect. Good retrieval does not guarantee a good answer, and bad retrieval does not guarantee a bad one. The cookbook says: 'A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.' That is why it insists on collecting both kinds of metric. If you only look at the final text, you can't tell whether to fix the retriever or the generation step.
Checkpoint 1 of 5· Check yourself
Your team only reviews final answers and they look acceptable. A colleague proposes adding retrieval metrics too. Which statement best supports that proposal?
The docs say response and retrieval quality can move apart in both directions. Measuring both is the only way to diagnose which component needs fixing.
“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”Source: docs.databricks.com
Sources1
2.Naming the common issues and the judges that catch them
Each issue you might spot by eye has a built-in MLflow judge with a matching question. In MLflow, a judge is a scorer that uses an LLM to assess quality, so it can handle meaning rather than exact wording. The docs give an example: a judge can tell 'give me healthy food options' and 'food to keep me fit' are similar queries. The table maps common issues to the judges that detect them.
| Issue you are looking for | Built-in judge | Question the judge asks |
|---|---|---|
| Off-topic answer | RelevanceToQuery | Is the response directly relevant to the user's request? |
| Irrelevant context fed to the model | RetrievalRelevance | Is the retrieved context directly relevant to the user's request? |
| Harmful, offensive or toxic content | Safety | Is the content free from harmful, offensive, or toxic material? |
| Hallucination | RetrievalGroundedness | Is the response grounded in the information provided in the context? |
| Factually wrong answer | Correctness | Is the response correct as compared to the provided ground truth? |
| Breaks style or policy rules | Guidelines | Does the response meet specified natural language criteria? |
| Wrong or redundant tool use | ToolCallCorrectness / ToolCallEfficiency | Are the tool calls correct / efficient without redundancy? |
One judge in this table, Safety, covers the *safety* half of the objective for a single response: it checks for harmful material. Its conversation-level counterpart, ConversationalSafety, is covered below. The rest of the table covers *quality*. RelevanceToQuery and RetrievalGroundedness are often confused. An answer can be perfectly relevant and still invented, and it can be fully grounded and still not answer the question.
A note on names: the cookbook calls this dimension *groundedness*, while the built-in MLflow judge that implements it is called RetrievalGroundedness. They are the same check, so a question that says 'the Groundedness judge' is pointing at RetrievalGroundedness.
Safety is broader than toxicity. The Safety judge asks about harmful, offensive or toxic content. Other safety problems are handled elsewhere. Databricks AI Gateway has built-in guardrails for unsafe content, jailbreak and hallucination. The Jailbreak guardrail 'denies prompt-injection and jailbreak attempts (requests only)', so it acts on the input side. Privacy rules can be written as Guidelines judges: the docs' customer service example has the rule that the response 'must never ask for full credit card numbers, SSN, or passwords'.
Checkpoint 2 of 5· Exam question
A team deployed a RAG chatbot that answers questions about internal HR policy. During qualitative review, an evaluator notices the chatbot states "employees receive 25 vacation days annually," but the retrieved context chunks for that response only mention "20 vacation days" and never state 25 anywhere. Which approach is best suited to catch this specific quality issue automatically?
Correct answer: A — Run the Groundedness judge, which flags response claims that are not supported by the retrieved context as hallucinated content
- A. The Groundedness judge is reference-free and specifically checks whether claims in a response are supported by the retrieved context, so a fabricated number like 25 days that never appears in the source chunks would be flagged as an ungrounded hallucination.
- B. Correctness requires a labeled ground-truth expected answer to compare against, which this scenario does not describe; it also measures overall answer accuracy rather than specifically checking claims against retrieved context.
- C. RelevanceToQuery only measures whether the response addresses the topic and intent of the question; a response can be perfectly on-topic (about vacation days) while still stating a fabricated number, so this judge would not catch the error.
- D. ToolCallCorrectness evaluates whether a tool was called with the right arguments, which is unrelated to whether the final text response accurately reflects the content that was retrieved.
Checkpoint 3 of 5· Check yourself
A reviewer flags responses that are on-topic and grounded but sometimes include insulting language toward the user. Which built-in judge targets this issue?
Insulting or toxic language is a safety issue, whatever the relevance or grounding. The Safety judge asks whether content is free from harmful, offensive or toxic material.
“Is the content free from harmful, offensive, or toxic material?”Source: docs.databricks.com
3.Which assessments need a reference answer
Some issues can be judged from the request and response alone. Others need you to know the right answer first. The cookbook draws the line this way: some judges, such as answer correctness, compare human-labeled ground truth against the app's output. Others don't need it: 'Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.'
This matters in practice. You can run RelevanceToQuery, Safety, RetrievalGroundedness and Guidelines on any traffic, including unlabeled production requests. Correctness, RetrievalSufficiency and ToolCallCorrectness need an expectations field. ExpectationsGuidelines is a special case: it doesn't require ground truth, but it does need per-example guidelines stored in expectations. In the cookbook's metric list, document_recall needs ground truth and is computed deterministically, and chunk_relevance/precision uses an LLM judge with no ground truth.
Checkpoint 4 of 5· Match them up
Match each judge to what it needs in order to run
Tap a term, then the definition that fits it.
Watch ExpectationsGuidelines: it reads the expectations field, but that field holds guidelines, not a ground-truth answer. Of these four judges, only Correctness needs a ground-truth answer.
“Is the response correct as compared to the provided ground truth?”Source: docs.databricks.com
Sources1
4.Issues that only appear across a conversation
Reviewing one turn at a time misses some problems. Maybe the assistant forgets what the user said three turns ago. Maybe it slowly drifts out of its assigned persona, or the user grows more frustrated with each reply. MLflow's multi-turn judges take a whole session instead of a single request and response: 'These judges analyze the complete conversation history to assess quality patterns that emerge over multiple interactions.'
| Issue | Multi-turn judge |
|---|---|
| User questions left unanswered | ConversationCompleteness |
| User frustration, and whether it was resolved | UserFrustration |
| Forgetting earlier information | KnowledgeRetention |
| Drifting out of the assigned role | ConversationalRoleAdherence |
| Harmful content anywhere in the session | ConversationalSafety |
| Breaking guidelines over the session | ConversationalGuidelines |
| Inefficient or inappropriate tool use across the session | ConversationalToolCallEfficiency |
Checkpoint 5 of 5· Check yourself
A support bot is told to act as a billing assistant. Over a long chat, it starts giving medical advice. Each reply looks harmless on its own. Which judge targets this?
Drifting from an assigned role is a pattern across the session. ConversationalRoleAdherence assesses whether the assistant keeps its role throughout the conversation.
“Does the assistant maintain its assigned role throughout the conversation?”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Detecting hallucination needs a labeled ground-truth answer for every request.Why is that wrong?
Groundedness checks the response against the retrieved context, not against a reference answer. It doesn't need human-labeled ground truth.
Covered in Which assessments need a reference answer
2.If the final answers look good, retrieval must be working, so retrieval metrics are unnecessary.Why is that wrong?
Good responses can come from faulty retrievals, and poor responses from correct context, so you must measure both to find the real cause.
Covered in Where response problems come from
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“how safe the response was (in other words, no toxicity)”
↩︎ Where response problems come from“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
↩︎ Where response problems come from“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
↩︎ Which assessments need a reference answer“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
↩︎ Exam trap 1“It is very important to collect both response and retrieval metrics.”
↩︎ Exam trap 2 - 2.
“a judge can understand that give me healthy food options and food to keep me fit are similar queries”
↩︎ Naming the common issues and the judges that catch them - 3.
“Is the response grounded in the information provided in the context? Is the agent hallucinating?”
↩︎ Naming the common issues and the judges that catch them“These judges analyze the complete conversation history to assess quality patterns that emerge over multiple interactions.”
↩︎ Issues that only appear across a conversation“use Databricks-hosted LLMs to evaluate common quality dimensions of your agent such as relevance, safety, groundedness, and correctness”
↩︎ Key concept“Is the content free from harmful, offensive, or toxic material?”
↩︎ Checkpoint“Is the response correct as compared to the provided ground truth?”
↩︎ Checkpoint“Does the assistant maintain its assigned role throughout the conversation?”
↩︎ Checkpoint - 4.
“denies prompt-injection and jailbreak attempts (requests only)”
↩︎ Naming the common issues and the judges that catch them - 5.https://docs.databricks.com/aws/en/mlflow3/genai/eval-monitor/concepts/judges/guidelinesOfficial docs
“The response must never ask for full credit card numbers, SSN, or passwords”
↩︎ Naming the common issues and the judges that catch them