CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 47/56

    LLM evaluation metrics for model selection: quality, cost and latency

    Select an LLM choice (size and architecture) based on a set of quantitative evaluation metrics

    10 min read
    1.79% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Name the retrieval, response, cost and latency metrics Databricks recommends for comparing candidate LLMs
    • Tell which judges need human-labeled ground truth and which do not
    • Build an evaluation set and run mlflow.genai.evaluate() so candidate models are compared on the same data and scorers

    Key concept

    Quality–cost–latency balance — Choosing an LLM means finding the candidate that meets your quality bar at an acceptable cost and latency. Measuring those dimensions with the same metrics for every candidate is what makes the comparison quantitative rather than a matter of taste.

    1.The dimensions you compare candidate models on

    Databricks Foundation Model APIs are built so you can efficiently compare LLMs to find the best candidate for your use case, or swap a production model for one that performs better. To make that comparison you need numbers, not impressions. For a RAG application the Databricks guidance splits those numbers into three groups. Retrieval quality measures whether the app found relevant supporting data. Response quality measures whether the answer is accurate against the ground truth, grounded in the retrieved context (or hallucinated), and safe (no toxicity). System performance covers cost and latency, such as overall latency and token consumption.

    There are two ways to compute these metrics. Deterministic measurement covers cost and latency, which come straight from the application's outputs, and some retrieval metrics when your evaluation set lists the documents that contain each answer. LLM judge-based measurement uses a separate LLM to grade retrieval and response quality. The table below lists the metrics Databricks recommends. Put cost and latency next to the quality scores when you compare candidates, and re-measure all of them whenever you change the model.

    Recommended metrics for comparing RAG application configurations
    DimensionMetricMeasured byNeeds ground truth?
    Retrievalchunk_relevance/precisionLLM judgeNo
    Retrievaldocument_recallDeterministicYes
    Retrievalcontext_sufficiencyLLM judgeYes
    ResponsecorrectnessLLM judgeYes
    Responserelevance_to_queryLLM judgeNo
    ResponsegroundednessLLM judgeNo
    ResponsesafetyLLM judgeNo
    Costtotal_token_count, total_input_token_count, total_output_token_countDeterministicNo
    Latencylatency_secondsDeterministicNo

    Model size is a selection variable, and Databricks guidance is to start small and scale up as needed. Smaller open-source models can often give satisfactory results at lower cost and with faster inference, while a larger model is worth it only if the smaller one struggles on your queries. Then monitor the effect of each change on response quality, latency and cost. Measure latency rather than guessing it. Databricks defines total latency as time to first token plus time per output token multiplied by the number of tokens generated, so output length matters, and latency also changes with the number of concurrent requests. Databricks provides a benchmarking notebook to load-test an LLM endpoint for this.

    Checkpoint 1 of 6· Check yourself

    Which pair of metrics can be computed deterministically, without an LLM judge?

    Checkpoint 2 of 6· Exam question

    A team is building a customer support chatbot that must respond within 300 milliseconds to feel conversational, and offline evaluation shows a 70-billion-parameter model and a 8-billion-parameter model both clear the required 85% task-accuracy threshold on the held-out evaluation set. Based on these quantitative results, which model choice best fits the deployment requirement?

    Sources1234

    2.Judges that need ground truth and judges that don't

    The last column of the metrics table matters when you plan a comparison. Some judges compare the app's output with a human-labeled answer, so they only work if your evaluation data includes one. Others grade the output against the request or the retrieved context, so they can score any input. The metrics table above uses Agent Evaluation metric names. MLflow has its own built-in judges, listed below, with similar but differently named checks (for example RetrievalRelevance and RetrievalGroundedness). In MLflow, built-in judges are predefined scorers that use Databricks-hosted LLMs. Each one takes inputs and outputs, and the ground-truth judges also take expectations.

    Built-in MLflow judges and whether they require ground truth
    JudgeArgumentsRequires ground truthWhat it evaluates
    RelevanceToQueryinputs, outputsNoIs the response directly relevant to the user's request?
    RetrievalRelevanceinputs, outputsNoIs the retrieved context directly relevant to the user's request?
    Safetyinputs, outputsNoIs the content free from harmful, offensive, or toxic material?
    RetrievalGroundednessinputs, outputsNoIs the response grounded in the information provided in the context? Is the agent hallucinating?
    Correctnessinputs, outputs, expectationsYesIs the response correct as compared to the provided ground truth?
    RetrievalSufficiencyinputs, outputs, expectationsYesDoes the context provide all necessary information to generate a response that includes the ground truth facts?
    ToolCallCorrectnessinputs, outputs, expectationsYesAre the tool calls and arguments correct for the user query?

    In practice you don't have to wait for a fully labeled dataset before comparing models. Agent Evaluation can score quality without ground truth, and once ground truth is available it adds metrics such as answer correctness. A sensible order is to compare candidates first on groundedness, relevance, safety, cost and latency, then add correctness as labels arrive. If no built-in judge fits your use case, scorers also come as custom LLM judges, code-based scorers for deterministic checks such as exact matching or format validation, and third-party scorers.

    Checkpoint 3 of 6· Match them up

    Match each judge to what it needs in order to run

    Tap a term, then the definition that fits it.

    Checkpoint 4 of 6· Exam question

    A financial services company is evaluating three candidate LLMs to summarize long regulatory filings that average 25,000 tokens per document. During evaluation, one candidate model has a maximum context window of 8,000 tokens, while the other two support 32,000 and 128,000 tokens respectively, and all three otherwise score similarly on summarization quality metrics. What should the team conclude from this quantitative comparison?

    Sources56

    3.A fair comparison: one evaluation set, one harness

    Metrics only let you compare models if every candidate is scored on the same inputs. Databricks recommends a human-labeled evaluation set: a curated, representative set of queries with ground-truth answers and, optionally, the supporting documents that should be retrieved. It should be representative of production traffic, challenging (including adversarial prompts such as prompt-injection attempts), and updated over time. Aim for at least 30 questions, ideally 100–200. To avoid overfitting while you try many configurations, split the set into training (~70%), test (~20%) and validation (~10%) portions.

    Checkpoint 5 of 6· Put it in order

    You are comparing six candidate LLM configurations. Put the evaluation-set splits in the order you use them.

    1. 1.Run every experiment on the training split (~70%) to find the highest-potential candidates
    2. 2.Run a final check on the validation split (~10%) before deploying to production
    3. 3.Evaluate the highest-performing experiments on the test split (~20%)

    The harness that runs the comparison is mlflow.genai.evaluate(). It runs your app over the evaluation data, applies the scorers you choose, and returns an EvaluationResult. Databricks lists validating prompt or model changes across app versions as one of its main uses. To compare two LLMs, keep data and scorers the same and change only the app wrapper, using the optional model_id to track which version produced which results.

    Signature of mlflow.genai.evaluate(): test data, scorers, the app wrapper, and optional version trackingpython
    def mlflow.genai.evaluate(
        data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset],  # Test data.
        scorers: list[mlflow.genai.scorers.Scorer],  # Quality metrics, built-in or custom.
        predict_fn: Optional[Callable[..., Any]] = None,  # App wrapper. Used for direct evaluation only.
        model_id: Optional[str] = None,  # Optional version tracking.
    ) -> mlflow.models.evaluation.base.EvaluationResult:

    Checkpoint 6 of 6· Fill the gap

    In direct evaluation, MLflow calls your app itself. Which parameter passes the app to the harness?

    results = mlflow.genai.evaluate(
        data=[
            {"inputs": {"question": "What is MLflow?"}},
            {"inputs": {"question": "How do I get started?"}}
        ],
         ? =my_chatbot_app,
        scorers=[RelevanceToQuery(), Safety()]
    )

    Sources57

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A smaller model is always faster, so you can pick it on parameter count without measuring latency.Why is that wrong?

      Databricks says to measure latency: it depends on time to first token and on how many output tokens are generated, and it changes with concurrent load. Smaller models are a good place to start, but you confirm with metrics on your own workload.

      Covered in The dimensions you compare candidate models on

    2. 2.You can't compare candidate LLMs until you have a fully ground-truth-labeled evaluation set.Why is that wrong?

      Agent Evaluation can assess quality without ground truth, using reference-free judges such as groundedness. Ground truth only adds metrics such as answer correctness.

      Covered in Judges that need ground truth and judges that don't

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Efficiently compare LLMs to see the best candidate for your use case, or swap a production model with a better performing one.”
      ↩︎ The dimensions you compare candidate models on
    2. 2.
      “Cost and latency metrics can be computed deterministically based on the application's outputs.”
      ↩︎ The dimensions you compare candidate models on
      “Only by measuring both components can we accurately diagnose and address issues in the application.”
      ↩︎ Prediction
      “Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
      ↩︎ Checkpoint
    3. 3.
      “Monitor the impact of changing models on key metrics such as response quality, latency, and cost”
      ↩︎ The dimensions you compare candidate models on
      “Selecting the most appropriate model (and model parameters) for your application to optimize/balance performance, latency, and cost.”
      ↩︎ Key concept
    4. 4.
      “Latency = TTFT + (TPOT) * (the number of tokens to be generated)”
      ↩︎ The dimensions you compare candidate models on
      “The number of output tokens dominates overall response latency.”
      ↩︎ Exam trap 1
    5. 5.
      “Agent Evaluation can assess your chain's quality without ground truth, although, if ground truth is available, it computes additional metrics such as answer correctness.”
      ↩︎ Judges that need ground truth and judges that don't
      “Databricks recommends at least 30 questions in your evaluation set, and ideally 100 - 200.”
      ↩︎ A fair comparison: one evaluation set, one harness
      “Agent Evaluation can assess your chain's quality without ground truth”
      ↩︎ Exam trap 2
      “Validation set: ~10% of the questions. Used for a final validation check before deploying an experiment to production.”
      ↩︎ Checkpoint
    6. 6.
      “Built-in LLM judges are predefined scorers that use Databricks-hosted LLMs to evaluate common quality dimensions of your agent”
      ↩︎ Judges that need ground truth and judges that don't

    Continue to page 2 of 2

    Choosing LLM size and serving setup from evaluation and benchmark results

    Spotted a mistake, or was something unclear? Tell us.