CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 53/56

    Ground Truth in RAG Metrics and Custom make_judge() Judges

    Identify evaluation judges that require ground truth

    8 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Classify RAG retrieval and response metrics by whether they need ground truth, including the deterministic document_recall metric
    • Explain why recall needs ground truth and precision does not
    • Use the {{ expectations }} template variable to make a custom judge depend on ground truth, or leave it out to keep the judge reference-free
    • Plan an evaluation that starts without labels and adds ground-truth judges later

    1.RAG metrics: ground truth is not only about LLM judges

    The Databricks RAG cookbook groups its recommended metrics into retrieval, response, cost and latency. It measures each one either deterministically or with an LLM judge. The ground-truth question cuts across both methods. The cookbook puts the LLM-judge split plainly: some judges, such as answer correctness, compare the human-labelled ground truth with the app outputs, and others, such as groundedness, do not need it.

    Recommended RAG metrics: how each is measured and whether it needs ground truth
    DimensionMetricMeasured byNeeds ground truth?
    Retrievalchunk_relevance/precisionLLM judgeNo
    Retrievaldocument_recallDeterministicYes
    Retrievalcontext_sufficiencyLLM JudgeYes
    ResponsecorrectnessLLM judgeYes
    Responserelevance_to_queryLLM judgeNo
    ResponsegroundednessLLM judgeNo
    ResponsesafetyLLM judgeNo
    Costtotal_token_count, total_input_token_count, total_output_token_countDeterministicNo
    Latencylatency_secondsDeterministicNo

    The retrieval pair shows the logic most clearly. Precision asks: of the chunks I retrieved, what share are relevant to the query? An LLM judge can decide relevance chunk by chunk, so you never need to know the full set of relevant documents. Recall asks: of all the documents known to be relevant, what share did I retrieve? That question has no answer unless someone has listed every relevant document beforehand. In the cookbook's example, two of the three retrieved results were relevant, so precision was 0.66. The retrieved results covered two of the four relevant documents, so recall was 0.5. In the MLflow dataset schema, this list is the expected_retrieved_context key, which the document_recall scorer reads.

    The same pattern holds for the response metrics. Correctness and context sufficiency need ground truth; relevance, groundedness and safety do not. These match the Yes/No answers that the built-in judge table gives Correctness, RetrievalSufficiency, RelevanceToQuery, RetrievalGroundedness and Safety. Cost and latency come straight from the app's execution, so they never need labels.

    Checkpoint 1 of 4· Check yourself

    Your evaluation set has questions only, with no labelled documents. Which retrieval metric can you still compute?

    Sources12

    2.Custom judges: you decide whether ground truth is required

    When the built-in judges don't fit, make_judge() lets you write the evaluation criteria in natural language. The instructions can reference only four template variables: {{ inputs }}, {{ outputs }}, {{ expectations }} and {{ trace }}. Custom names such as {{ question }} cause validation errors. Of the four, {{ expectations }} is the one documented as "Ground truths or expected outcomes". So whether a custom judge requires ground truth is not a fixed property of the judge type. It depends on whether your instructions reference {{ expectations }}.

    A field-based custom judge that requires ground truth: it compares {{ outputs }} with {{ expectations }}python
    from mlflow.genai.judges import make_judge
    from typing import Literal
    
    # Field-based judge: compare the response to the expected answer
    correctness_judge = make_judge(
        name="answer_correctness",
        instructions=(
            "Given the request in {{ inputs }}, decide whether the response in "
            "{{ outputs }} matches the expected answer in {{ expectations }}.\n\n"
            "Return 'correct' only if the response conveys the same facts as the "
            "expected answer. Otherwise, return 'incorrect'."
        ),
        feedback_value_type=Literal["correct", "incorrect"],
    )

    The docs give the reference-free option directly: leave out {{ expectations }} when you don't have ground truth, for example to score tone or format from the inputs and outputs alone. The instructions must include at least one template variable, but they don't have to use all four. Trace-based judges work the same way. The documented tool-usage validator inspects only {{ trace }} to check that the agent picked appropriate tools with correct parameters, so it needs no labels. Note the contrast with the built-in ToolCallCorrectness judge, which compares against expectations.

    Checkpoint 2 of 4· Check yourself

    You need a custom judge that scores the tone of support replies, and you have no labelled answers. Which template should the instructions use?

    Checkpoint 3 of 4· Exam question

    An agent invokes external tools such as a calculator and a search API. The evaluation team has a dataset specifying, for each test case, the exact tool names and arguments the agent should have called. They want a judge that flags cases where the agent's actual tool invocations differ from this expected sequence. Which judge requires this kind of ground truth?

    Sources3

    3.Starting without labels, adding ground truth later

    Knowing which judges need ground truth tells you what you can measure today. Databricks recommends a human-labelled evaluation set, but admits that curating labels takes time. Its advice is to start with an evaluation set that contains only questions and add ground-truth responses over time. Agent Evaluation can assess quality without ground truth, and it computes extra metrics such as answer correctness once ground truth exists.

    In practice, your first runs use reference-free judges on unlabelled data. The harness example below scores pre-computed outputs that have no expectations at all:

    Scoring pre-computed outputs with reference-free judges only, with no expectations in the datapython
    # Evaluate pre-computed outputs
    evaluation = mlflow.genai.evaluate(
        data=results_data,
        scorers=[Safety(), RelevanceToQuery()]
    )

    As experts add expected_facts, expected_response or expected_retrieved_context to rows, you can add Correctness, RetrievalSufficiency and document_recall to the scorer list. Those additions are the metrics that tell you whether the answer was actually right, not just relevant and safe. The cookbook warns that you need both kinds of metric. A RAG app can answer poorly despite retrieving the correct context, and it can answer well despite faulty retrieval.

    Checkpoint 4 of 4· Check yourself

    A team's evaluation set has 50 questions and no labelled answers. What does the Databricks guidance say?

    Sources41

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Deterministic metrics never need ground truth; only LLM judges do.Why is that wrong?

      document_recall is deterministic, and it still needs the labelled list of relevant documents to measure what share were retrieved.

      Covered in RAG metrics: ground truth is not only about LLM judges

    2. 2.Every custom make_judge() judge needs ground truth because it is a custom evaluation.Why is that wrong?

      A custom judge needs ground truth only if its instructions reference {{ expectations }}. You must use at least one of the four variables, not all of them.

      Covered in Custom judges: you decide whether ground truth is required

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Some LLM judges, such as answer correctness, compare the human-labeled ground truth vs. the app outputs.”
      ↩︎ RAG metrics: ground truth is not only about LLM judges
      “Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
      ↩︎ RAG metrics: ground truth is not only about LLM judges
      “A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
      ↩︎ Starting without labels, adding ground truth later
      “Computing recall requires your ground-truth to contain all relevant items.”
      ↩︎ Exam trap 1
      “document_recall | What % of the ground truth documents are represented in the retrieved chunks? | Deterministic | Yes”
      ↩︎ Prediction
      “Computing precision does not require knowing all relevant items.”
      ↩︎ Checkpoint
    2. 2.
      “expected_retrieved_context | document_recall scorer | Documents that should be retrieved”
      ↩︎ RAG metrics: ground truth is not only about LLM judges
    3. 3.
      “{{ expectations }} - Ground truths or expected outcomes”
      ↩︎ Custom judges: you decide whether ground truth is required
      “Analyze the {{ trace }} to verify correct tool usage.”
      ↩︎ Custom judges: you decide whether ground truth is required
      “Your instructions must include at least one template variable, but you don't need to use all of them.”
      ↩︎ Exam trap 2
      “Omit {{ expectations }} when you don't have ground truth, for example to score tone or format from the inputs and outputs alone.”
      ↩︎ Checkpoint
    4. 4.
      “You can get started by creating an evaluation set that only includes questions, and add the ground truth responses over time.”
      ↩︎ Starting without labels, adding ground truth later
      “Agent Evaluation can assess your chain's quality without ground truth, although, if ground truth is available, it computes additional metrics such as answer correctness.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.