What you will be able to do
- Apply the start-small, scale-up-on-evidence approach to choosing an LLM's size
- Decide when to use different or fine-tuned smaller models for different steps of a chain
- Read TTFT, TPOT, latency and throughput benchmarks and match them to real-time or batch requirements
1.Model size: start small and scale up on evidence
Choosing a model's size is a loop driven by your metrics. Databricks advises against jumping straight to the most powerful models. Smaller open-source alternatives such as Llama 3 can often give satisfactory results at lower cost and with faster inference, especially for tasks that don't need complex reasoning or broad world knowledge. As you refine the chain, keep checking where the chosen model falls short. If it struggles with certain query types, or its answers aren't detailed or accurate enough, scale up to a more capable model. Each time you swap, re-measure response quality, latency and cost so you can see whether the trade-off still fits your use case.
Two practical points come with every swap. First, prompts often don't carry over between models: a prompt tuned for one model may work worse on another, so expect to adapt it, and give each candidate a fair prompt before comparing their scores. Second, parameters such as temperature (randomness) and max_tokens (response length) affect both quality and output length, so tune and evaluate them as part of the comparison.
Checkpoint 1 of 3· Check yourself
A small open model scores well on groundedness and safety but fails correctness on multi-step reasoning questions in your evaluation set. What does the guidance recommend?
Failing on specific query types is the signal to scale up, and the swap should be checked against response quality, latency and cost. Removing hard questions defeats the purpose of a challenging evaluation set.
“Monitor the impact of changing models on key metrics such as response quality, latency, and cost”Source: docs.databricks.com
Sources1
2.Architecture: different models for different steps
A RAG chain often calls an LLM more than once, for example to understand the query before the main generation step. Each call can use its own model. A large, general-purpose model may be overkill for a narrow job like working out the user's intent, so a cheaper model there lowers cost and latency without hurting the final answer. You can go further: fine-tuning smaller models for specific sub-tasks such as query understanding can improve overall performance, reduce latency and lower inference costs compared with one large model for everything. If the domain uses terminology the base model doesn't know well, continued pre-training on domain data is another option.
To compare candidates quantitatively, run each one on the same evaluation set and put the aggregate metrics side by side. Databricks groups them into response quality (for example correctness and groundedness), and system performance, meaning cost and latency. Cost and latency are computed deterministically from the app's outputs: token counts stand in for cost and latency_seconds for speed. Quality metrics such as correctness need ground truth, while groundedness does not. Then decide against your requirements. One approach shown in the Databricks examples is a weighted composite of metrics, such as weighting correctness 70% and another metric 30%.
| Candidate | correctness | groundedness | latency_seconds | total_token_count |
|---|---|---|---|---|
| Small model | 0.78 | 0.90 | 1.2 | 900 |
| Mid-size model | 0.83 | 0.91 | 2.1 | 1100 |
| Large model | 0.86 | 0.92 | 3.4 | 1500 |
Reading the table: the small model is fast and cheap but misses the correctness bar, so the evidence says scale up. The large model is the most accurate but breaks the latency budget and uses the most tokens. The mid-size model is the only one that meets both requirements, so it is the choice. The numbers are invented for illustration. The method is the point: set thresholds from your use case, then pick the smallest candidate that clears them, and re-measure whenever you swap.
Checkpoint 2 of 3· Check yourself
Your chain uses one large frontier model for both intent classification and final answer generation. Token cost is over budget. Which change does the guidance support?
Intent detection is a narrow task where an expensive general-purpose model is overkill. A smaller or fine-tuned model there cuts cost and latency.
“More expensive, general-purpose models may be overkill for tasks like determining the intent of a user query.”Source: docs.databricks.com
3.Serving metrics: TTFT, TPOT, latency and throughput
Quality scores tell you whether a model is good enough. Serving benchmarks tell you whether it is fast enough at the load you expect. LLM inference happens in two steps: prefill processes the prompt's tokens in parallel, then decoding generates output one token at a time. Databricks breaks endpoint performance into two sub-metrics. Time to first token (TTFT) is how quickly users start seeing output. It matters most in real-time use and much less offline. Time per output token (TPOT) is how fast the model feels to each user. For example, 100 ms per token is 10 tokens per second.
From those two numbers: Latency = TTFT + TPOT × (number of tokens to be generated), and throughput is output tokens per second across all concurrent requests. Because generated tokens multiply TPOT, the number of output tokens dominates response latency. Input tokens mainly affect how much memory a request needs. That is why max_tokens and model verbosity show up directly in your latency numbers.
| Use case | Example | What to prioritise | Concurrency strategy |
|---|---|---|---|
| Low latency | Real-time applications that need immediate responses | Low TTFT and total latency | Send fewer concurrent requests to keep latency low |
| High throughput | Batch inference and other non-user-facing tasks | Output tokens per second | Saturate the endpoint with many concurrent requests |
Latency and throughput trade off against each other. As concurrency rises, throughput usually climbs and so does latency, until throughput levels off. In the Databricks example benchmark notebook it plateaus at about 8,000 tokens per second. That figure is what the notebook shows for its endpoint, not a general constant: the plateau occurs because the endpoint's provisioned throughput limits how many workers and parallel requests it can handle. Beyond that point, extra requests just wait in the queue. Databricks' advice is to maximise throughput within your latency budget, and it recommends provisioned throughput for production workloads. So when comparing candidate sizes, benchmark each one at your expected concurrency and check it stays within the latency budget.
Checkpoint 3 of 3· Check yourself
Benchmarks for a chat assistant show total latency above budget. The prompts are short, but answers average several hundred tokens. Which lever most directly reduces latency?
Total latency is TTFT plus TPOT times the generated tokens, so output length dominates. Raising concurrency increases throughput but tends to increase latency.
“The number of output tokens dominates overall response latency.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Start with the most capable model, then downsize only if cost becomes a problem.Why is that wrong?
Databricks recommends starting with smaller, lightweight models and scaling up only when evaluation shows the model struggles.
2.Long input prompts are the main driver of response latency.Why is that wrong?
Output tokens dominate total latency, because each generated token adds TPOT. Input tokens mainly affect the memory needed to process the request.
Covered in Serving metrics: TTFT, TPOT, latency and throughput
3.A chain should use one model everywhere so its quality is consistent.Why is that wrong?
Databricks suggests using different models for different steps. Smaller or fine-tuned models handle narrow sub-tasks such as query understanding at lower cost and latency.
Covered in Architecture: different models for different steps
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“consider scaling up to a more capable model”
↩︎ Model size: start small and scale up on evidence“prompts often do not transfer seamlessly across different language models.”
↩︎ Model size: start small and scale up on evidence“consider using different models for different steps.”
↩︎ Architecture: different models for different steps“smaller open-source alternatives like Llama 3 can provide satisfactory results at a lower cost and with faster inference times.”
↩︎ Exam trap 1“consider fine-tuning smaller models for specific sub-tasks within your RAG chain, such as query understanding.”
↩︎ Exam trap 3“it's often more efficient to start with smaller, more lightweight models.”
↩︎ Prediction“Monitor the impact of changing models on key metrics such as response quality, latency, and cost”
↩︎ Checkpoint“More expensive, general-purpose models may be overkill for tasks like determining the intent of a user query.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“Overall latency and token consumption are examples of chain performance metrics.”
↩︎ Architecture: different models for different steps“Cost and latency metrics can be computed deterministically based on the application's outputs.”
↩︎ Architecture: different models for different steps - 3.https://docs.databricks.com/aws/en/mlflow3/genai/prompt-version-mgmt/prompt-registry/evaluate-promptsOfficial docs
“Weight correctness more heavily (70%) than compliance (30%)”
↩︎ Architecture: different models for different steps - 4.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/prov-throughput-run-benchmarkOfficial docs
“Latency = TTFT + (TPOT) * (the number of tokens to be generated)”
↩︎ Serving metrics: TTFT, TPOT, latency and throughput“Most production applications have a latency budget, and Databricks recommends you maximize throughput given that latency budget.”
↩︎ Serving metrics: TTFT, TPOT, latency and throughput“This plateau occurs because the provisioned throughput for the endpoint limits the number of workers and parallel requests that can be made.”
↩︎ Serving metrics: TTFT, TPOT, latency and throughput“The number of input tokens has a substantial impact on the required memory to process requests.”
↩︎ Exam trap 2“The number of output tokens dominates overall response latency.”
↩︎ Prediction - 5.
“Databricks recommends provisioned throughput for production workloads.”
↩︎ Serving metrics: TTFT, TPOT, latency and throughput