CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 16/56

    Quality and Safety Issues in GenAI Responses: What to Look For

    Qualitatively assess responses to identify common issues such as quality and safety

    9 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Separate retrieval problems from response problems when you review a RAG answer
    • Match common response issues (irrelevance, hallucination, unsafe content, incorrectness, guideline violations) to the MLflow built-in judge that detects each one
    • Tell which assessments need a ground-truth expectation and which do not
    • Recognise quality and safety issues that only show up across a whole conversation

    Key concept

    Quality dimension (as assessed by an LLM judge) — A named part of response quality, such as relevance, groundedness, safety or correctness, that you assess on its own. Each one has a question you can ask about a response, and an LLM judge can answer it at scale. If you break 'is this answer good?' into separate dimensions, you can say exactly what went wrong.

    1.Where response problems come from

    When you assess a response qualitatively, the first job is to find where the issue started. Databricks splits RAG application performance into three dimensions. Retrieval quality asks whether the app fetched relevant supporting data. Response quality asks how well the app answered the user: whether the answer matches the ground truth, whether it is grounded in the retrieved context (or hallucinated), and whether it is safe. The source puts safety as 'how safe the response was (in other words, no toxicity)'. System performance covers cost and latency, such as token consumption and overall latency.

    The first two are tied together, but not in the way people expect. Good retrieval does not guarantee a good answer, and bad retrieval does not guarantee a bad one. The cookbook says: 'A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.' That is why it insists on collecting both kinds of metric. If you only look at the final text, you can't tell whether to fix the retriever or the generation step.

    Checkpoint 1 of 5· Check yourself

    Your team only reviews final answers and they look acceptable. A colleague proposes adding retrieval metrics too. Which statement best supports that proposal?

    Sources1

    2.Naming the common issues and the judges that catch them

    Each issue you might spot by eye has a built-in MLflow judge with a matching question. In MLflow, a judge is a scorer that uses an LLM to assess quality, so it can handle meaning rather than exact wording. The docs give an example: a judge can tell 'give me healthy food options' and 'food to keep me fit' are similar queries. The table maps common issues to the judges that detect them.

    Common response issues and the built-in judge that assesses each
    Issue you are looking forBuilt-in judgeQuestion the judge asks
    Off-topic answerRelevanceToQueryIs the response directly relevant to the user's request?
    Irrelevant context fed to the modelRetrievalRelevanceIs the retrieved context directly relevant to the user's request?
    Harmful, offensive or toxic contentSafetyIs the content free from harmful, offensive, or toxic material?
    HallucinationRetrievalGroundednessIs the response grounded in the information provided in the context?
    Factually wrong answerCorrectnessIs the response correct as compared to the provided ground truth?
    Breaks style or policy rulesGuidelinesDoes the response meet specified natural language criteria?
    Wrong or redundant tool useToolCallCorrectness / ToolCallEfficiencyAre the tool calls correct / efficient without redundancy?

    One judge in this table, Safety, covers the *safety* half of the objective for a single response: it checks for harmful material. Its conversation-level counterpart, ConversationalSafety, is covered below. The rest of the table covers *quality*. RelevanceToQuery and RetrievalGroundedness are often confused. An answer can be perfectly relevant and still invented, and it can be fully grounded and still not answer the question.

    A note on names: the cookbook calls this dimension *groundedness*, while the built-in MLflow judge that implements it is called RetrievalGroundedness. They are the same check, so a question that says 'the Groundedness judge' is pointing at RetrievalGroundedness.

    Safety is broader than toxicity. The Safety judge asks about harmful, offensive or toxic content. Other safety problems are handled elsewhere. Databricks AI Gateway has built-in guardrails for unsafe content, jailbreak and hallucination. The Jailbreak guardrail 'denies prompt-injection and jailbreak attempts (requests only)', so it acts on the input side. Privacy rules can be written as Guidelines judges: the docs' customer service example has the rule that the response 'must never ask for full credit card numbers, SSN, or passwords'.

    Checkpoint 2 of 5· Exam question

    A team deployed a RAG chatbot that answers questions about internal HR policy. During qualitative review, an evaluator notices the chatbot states "employees receive 25 vacation days annually," but the retrieved context chunks for that response only mention "20 vacation days" and never state 25 anywhere. Which approach is best suited to catch this specific quality issue automatically?

    Checkpoint 3 of 5· Check yourself

    A reviewer flags responses that are on-topic and grounded but sometimes include insulting language toward the user. Which built-in judge targets this issue?

    Sources2345

    3.Which assessments need a reference answer

    Some issues can be judged from the request and response alone. Others need you to know the right answer first. The cookbook draws the line this way: some judges, such as answer correctness, compare human-labeled ground truth against the app's output. Others don't need it: 'Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.'

    This matters in practice. You can run RelevanceToQuery, Safety, RetrievalGroundedness and Guidelines on any traffic, including unlabeled production requests. Correctness, RetrievalSufficiency and ToolCallCorrectness need an expectations field. ExpectationsGuidelines is a special case: it doesn't require ground truth, but it does need per-example guidelines stored in expectations. In the cookbook's metric list, document_recall needs ground truth and is computed deterministically, and chunk_relevance/precision uses an LLM judge with no ground truth.

    Checkpoint 4 of 5· Match them up

    Match each judge to what it needs in order to run

    Tap a term, then the definition that fits it.

    Sources1

    4.Issues that only appear across a conversation

    Reviewing one turn at a time misses some problems. Maybe the assistant forgets what the user said three turns ago. Maybe it slowly drifts out of its assigned persona, or the user grows more frustrated with each reply. MLflow's multi-turn judges take a whole session instead of a single request and response: 'These judges analyze the complete conversation history to assess quality patterns that emerge over multiple interactions.'

    Conversation-level issues and their multi-turn judges
    IssueMulti-turn judge
    User questions left unansweredConversationCompleteness
    User frustration, and whether it was resolvedUserFrustration
    Forgetting earlier informationKnowledgeRetention
    Drifting out of the assigned roleConversationalRoleAdherence
    Harmful content anywhere in the sessionConversationalSafety
    Breaking guidelines over the sessionConversationalGuidelines
    Inefficient or inappropriate tool use across the sessionConversationalToolCallEfficiency

    Checkpoint 5 of 5· Check yourself

    A support bot is told to act as a billing assistant. Over a long chat, it starts giving medical advice. Each reply looks harmless on its own. Which judge targets this?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Detecting hallucination needs a labeled ground-truth answer for every request.Why is that wrong?

      Groundedness checks the response against the retrieved context, not against a reference answer. It doesn't need human-labeled ground truth.

      Covered in Which assessments need a reference answer

    2. 2.If the final answers look good, retrieval must be working, so retrieval metrics are unnecessary.Why is that wrong?

      Good responses can come from faulty retrievals, and poor responses from correct context, so you must measure both to find the real cause.

      Covered in Where response problems come from

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “how safe the response was (in other words, no toxicity)”
      ↩︎ Where response problems come from
      “A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
      ↩︎ Where response problems come from
      “Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
      ↩︎ Which assessments need a reference answer
      “Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
      ↩︎ Exam trap 1
      “It is very important to collect both response and retrieval metrics.”
      ↩︎ Exam trap 2
    2. 2.
      “a judge can understand that give me healthy food options and food to keep me fit are similar queries”
      ↩︎ Naming the common issues and the judges that catch them
    3. 3.
      “Is the response grounded in the information provided in the context? Is the agent hallucinating?”
      ↩︎ Naming the common issues and the judges that catch them
      “These judges analyze the complete conversation history to assess quality patterns that emerge over multiple interactions.”
      ↩︎ Issues that only appear across a conversation
      “use Databricks-hosted LLMs to evaluate common quality dimensions of your agent such as relevance, safety, groundedness, and correctness”
      ↩︎ Key concept
      “Is the content free from harmful, offensive, or toxic material?”
      ↩︎ Checkpoint
      “Is the response correct as compared to the provided ground truth?”
      ↩︎ Checkpoint
      “Does the assistant maintain its assigned role throughout the conversation?”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Reviewing GenAI Responses with MLflow Judges, Traces and Human Feedback

    Spotted a mistake, or was something unclear? Tell us.