What you will be able to do
- Name the retrieval, response, cost and latency metrics Databricks recommends for comparing candidate LLMs
- Tell which judges need human-labeled ground truth and which do not
- Build an evaluation set and run mlflow.genai.evaluate() so candidate models are compared on the same data and scorers
Key concept
Quality–cost–latency balance — Choosing an LLM means finding the candidate that meets your quality bar at an acceptable cost and latency. Measuring those dimensions with the same metrics for every candidate is what makes the comparison quantitative rather than a matter of taste.
1.The dimensions you compare candidate models on
Databricks Foundation Model APIs are built so you can efficiently compare LLMs to find the best candidate for your use case, or swap a production model for one that performs better. To make that comparison you need numbers, not impressions. For a RAG application the Databricks guidance splits those numbers into three groups. Retrieval quality measures whether the app found relevant supporting data. Response quality measures whether the answer is accurate against the ground truth, grounded in the retrieved context (or hallucinated), and safe (no toxicity). System performance covers cost and latency, such as overall latency and token consumption.
There are two ways to compute these metrics. Deterministic measurement covers cost and latency, which come straight from the application's outputs, and some retrieval metrics when your evaluation set lists the documents that contain each answer. LLM judge-based measurement uses a separate LLM to grade retrieval and response quality. The table below lists the metrics Databricks recommends. Put cost and latency next to the quality scores when you compare candidates, and re-measure all of them whenever you change the model.
| Dimension | Metric | Measured by | Needs ground truth? |
|---|---|---|---|
| Retrieval | chunk_relevance/precision | LLM judge | No |
| Retrieval | document_recall | Deterministic | Yes |
| Retrieval | context_sufficiency | LLM judge | Yes |
| Response | correctness | LLM judge | Yes |
| Response | relevance_to_query | LLM judge | No |
| Response | groundedness | LLM judge | No |
| Response | safety | LLM judge | No |
| Cost | total_token_count, total_input_token_count, total_output_token_count | Deterministic | No |
| Latency | latency_seconds | Deterministic | No |
Model size is a selection variable, and Databricks guidance is to start small and scale up as needed. Smaller open-source models can often give satisfactory results at lower cost and with faster inference, while a larger model is worth it only if the smaller one struggles on your queries. Then monitor the effect of each change on response quality, latency and cost. Measure latency rather than guessing it. Databricks defines total latency as time to first token plus time per output token multiplied by the number of tokens generated, so output length matters, and latency also changes with the number of concurrent requests. Databricks provides a benchmarking notebook to load-test an LLM endpoint for this.
Checkpoint 1 of 6· Check yourself
Which pair of metrics can be computed deterministically, without an LLM judge?
Cost (token counts) and latency are read directly from the app's outputs. The quality metrics in the other options are all graded by an LLM judge.
“Cost and latency metrics can be computed deterministically based on the application's outputs.”Source: docs.databricks.com
Checkpoint 2 of 6· Exam question
A team is building a customer support chatbot that must respond within 300 milliseconds to feel conversational, and offline evaluation shows a 70-billion-parameter model and a 8-billion-parameter model both clear the required 85% task-accuracy threshold on the held-out evaluation set. Based on these quantitative results, which model choice best fits the deployment requirement?
Correct answer: A — The 8-billion-parameter model, because it meets the accuracy threshold while its smaller size gives it lower inference latency for the response-time requirement
- A. This is correct: once both models clear the accuracy bar, the deciding quantitative factor is latency, and the smaller model typically has fewer layers and less compute per token, giving it lower inference latency to hit a tight response-time target.
- B. This is incorrect because both models already cleared the same accuracy threshold in the evaluation, so the larger model's parameter count offers no measured accuracy advantage here, only added latency and cost.
- C. This is incorrect because clearing an accuracy threshold does not exempt a model from further evaluation on latency, cost, or other production requirements; model selection should still weigh the response-time constraint.
- D. This is incorrect because Databricks provisioned throughput serving does not impose a minimum parameter count; eligibility depends on whether a specific model architecture is optimized for throughput-based serving, not its size.
2.Judges that need ground truth and judges that don't
The last column of the metrics table matters when you plan a comparison. Some judges compare the app's output with a human-labeled answer, so they only work if your evaluation data includes one. Others grade the output against the request or the retrieved context, so they can score any input. The metrics table above uses Agent Evaluation metric names. MLflow has its own built-in judges, listed below, with similar but differently named checks (for example RetrievalRelevance and RetrievalGroundedness). In MLflow, built-in judges are predefined scorers that use Databricks-hosted LLMs. Each one takes inputs and outputs, and the ground-truth judges also take expectations.
| Judge | Arguments | Requires ground truth | What it evaluates |
|---|---|---|---|
| RelevanceToQuery | inputs, outputs | No | Is the response directly relevant to the user's request? |
| RetrievalRelevance | inputs, outputs | No | Is the retrieved context directly relevant to the user's request? |
| Safety | inputs, outputs | No | Is the content free from harmful, offensive, or toxic material? |
| RetrievalGroundedness | inputs, outputs | No | Is the response grounded in the information provided in the context? Is the agent hallucinating? |
| Correctness | inputs, outputs, expectations | Yes | Is the response correct as compared to the provided ground truth? |
| RetrievalSufficiency | inputs, outputs, expectations | Yes | Does the context provide all necessary information to generate a response that includes the ground truth facts? |
| ToolCallCorrectness | inputs, outputs, expectations | Yes | Are the tool calls and arguments correct for the user query? |
In practice you don't have to wait for a fully labeled dataset before comparing models. Agent Evaluation can score quality without ground truth, and once ground truth is available it adds metrics such as answer correctness. A sensible order is to compare candidates first on groundedness, relevance, safety, cost and latency, then add correctness as labels arrive. If no built-in judge fits your use case, scorers also come as custom LLM judges, code-based scorers for deterministic checks such as exact matching or format validation, and third-party scorers.
Checkpoint 3 of 6· Match them up
Match each judge to what it needs in order to run
Tap a term, then the definition that fits it.
Correctness and RetrievalSufficiency compare against labeled expectations. Groundedness and safety judge the output on its own terms, so they need no labels.
“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”Source: docs.databricks.com
Checkpoint 4 of 6· Exam question
A financial services company is evaluating three candidate LLMs to summarize long regulatory filings that average 25,000 tokens per document. During evaluation, one candidate model has a maximum context window of 8,000 tokens, while the other two support 32,000 and 128,000 tokens respectively, and all three otherwise score similarly on summarization quality metrics. What should the team conclude from this quantitative comparison?
Correct answer: A — The model with an 8,000-token context window should be eliminated because it cannot ingest a full filing in a single pass, forcing lossy chunking that the other two candidates avoid
- A. This is correct: a maximum context window of 8,000 tokens cannot fit a 25,000-token filing, so that model would require chunking or truncation to process the document, which is a real, measurable limitation the other two candidates do not have at this document length.
- B. This is incorrect because context window size directly caps how much text a model can process in a single call; a document larger than the window cannot be summarized in one pass regardless of cost considerations.
- C. This is incorrect because both remaining candidates support the full 25,000-token document in a single call, so context window is no longer the differentiator between them, and dismissing the 128,000-token option on window size alone is not supported by the given data.
- D. This is incorrect because an 8,000-token window is too small to hold the full 25,000-token filing at all, so it cannot be preferred on the basis of the document fitting within it.
3.A fair comparison: one evaluation set, one harness
Metrics only let you compare models if every candidate is scored on the same inputs. Databricks recommends a human-labeled evaluation set: a curated, representative set of queries with ground-truth answers and, optionally, the supporting documents that should be retrieved. It should be representative of production traffic, challenging (including adversarial prompts such as prompt-injection attempts), and updated over time. Aim for at least 30 questions, ideally 100–200. To avoid overfitting while you try many configurations, split the set into training (~70%), test (~20%) and validation (~10%) portions.
Checkpoint 5 of 6· Put it in order
You are comparing six candidate LLM configurations. Put the evaluation-set splits in the order you use them.
- 1.Run every experiment on the training split (~70%) to find the highest-potential candidates
- 2.Run a final check on the validation split (~10%) before deploying to production
- 3.Evaluate the highest-performing experiments on the test split (~20%)
The training split screens every experiment, the test split narrows the field, and the validation split is held back for one final check before deployment.
“Validation set: ~10% of the questions. Used for a final validation check before deploying an experiment to production.”Source: docs.databricks.com
The harness that runs the comparison is mlflow.genai.evaluate(). It runs your app over the evaluation data, applies the scorers you choose, and returns an EvaluationResult. Databricks lists validating prompt or model changes across app versions as one of its main uses. To compare two LLMs, keep data and scorers the same and change only the app wrapper, using the optional model_id to track which version produced which results.
def mlflow.genai.evaluate(
data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset], # Test data.
scorers: list[mlflow.genai.scorers.Scorer], # Quality metrics, built-in or custom.
predict_fn: Optional[Callable[..., Any]] = None, # App wrapper. Used for direct evaluation only.
model_id: Optional[str] = None, # Optional version tracking.
) -> mlflow.models.evaluation.base.EvaluationResult:Checkpoint 6 of 6· Fill the gap
In direct evaluation, MLflow calls your app itself. Which parameter passes the app to the harness?
results = mlflow.genai.evaluate(
data=[
{"inputs": {"question": "What is MLflow?"}},
{"inputs": {"question": "How do I get started?"}}
],
? =my_chatbot_app,
scorers=[RelevanceToQuery(), Safety()]
)predict_fn wraps the app's entry point, so MLflow can run it on each input and capture traces. model_id is only for version tracking.
Source: docs.databricks.comExam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A smaller model is always faster, so you can pick it on parameter count without measuring latency.Why is that wrong?
Databricks says to measure latency: it depends on time to first token and on how many output tokens are generated, and it changes with concurrent load. Smaller models are a good place to start, but you confirm with metrics on your own workload.
2.You can't compare candidate LLMs until you have a fully ground-truth-labeled evaluation set.Why is that wrong?
Agent Evaluation can assess quality without ground truth, using reference-free judges such as groundedness. Ground truth only adds metrics such as answer correctness.
Covered in Judges that need ground truth and judges that don't
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Efficiently compare LLMs to see the best candidate for your use case, or swap a production model with a better performing one.”
↩︎ The dimensions you compare candidate models on - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“Cost and latency metrics can be computed deterministically based on the application's outputs.”
↩︎ The dimensions you compare candidate models on“Only by measuring both components can we accurately diagnose and address issues in the application.”
↩︎ Prediction“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
↩︎ Checkpoint - 3.
“Monitor the impact of changing models on key metrics such as response quality, latency, and cost”
↩︎ The dimensions you compare candidate models on“Selecting the most appropriate model (and model parameters) for your application to optimize/balance performance, latency, and cost.”
↩︎ Key concept - 4.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/prov-throughput-run-benchmarkOfficial docs
“Latency = TTFT + (TPOT) * (the number of tokens to be generated)”
↩︎ The dimensions you compare candidate models on“The number of output tokens dominates overall response latency.”
↩︎ Exam trap 1 - 5.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-define-qualityOfficial docs
“Agent Evaluation can assess your chain's quality without ground truth, although, if ground truth is available, it computes additional metrics such as answer correctness.”
↩︎ Judges that need ground truth and judges that don't“Databricks recommends at least 30 questions in your evaluation set, and ideally 100 - 200.”
↩︎ A fair comparison: one evaluation set, one harness“Agent Evaluation can assess your chain's quality without ground truth”
↩︎ Exam trap 2“Validation set: ~10% of the questions. Used for a final validation check before deploying an experiment to production.”
↩︎ Checkpoint - 6.
“Built-in LLM judges are predefined scorers that use Databricks-hosted LLMs to evaluate common quality dimensions of your agent”
↩︎ Judges that need ground truth and judges that don't - 7.
“Validating prompt or model changes across app versions”
↩︎ A fair comparison: one evaluation set, one harness