What you will be able to do
- Explain why you would serve several models behind one Model Serving endpoint and split real-time traffic between them
- Read and write the served_entities and traffic_config.routes sections of an endpoint configuration
- Set an initial traffic split with the REST API, the MLflow Deployments SDK or the Serving UI
- Change an existing split with PUT /api/2.0/serving-endpoints/{name}/config without recreating the endpoint
- Send requests straight to one served model, skipping the traffic split, and know when that is useful
Key concept
Traffic configuration (traffic_config) — The part of a Model Serving endpoint's configuration that says what share of the endpoint's incoming requests each served model gets. You need it once an endpoint serves more than one model, and it is how real-time A/B tests and champion-versus-challenger comparisons are run.
1.Why split real-time traffic across models
Once a model is behind a Model Serving endpoint, every client calls one REST API. At some point you will train a better model, or think you have. There are two ways to check before switching all users to it. An offline comparison runs both models against a held-out data set and records the results with MLflow Tracking. An online comparison sends live requests to both models at once and compares how they behave in production. Databricks' MLOps guidance says real-time serving is where online comparisons, such as A/B tests or a gradual rollout, are most useful.
Model Serving lets you run an online comparison without a second endpoint or any client changes. One endpoint can host several served models, and you set what percentage of requests each one receives. A common setup keeps the current production model (the "champion") on most of the traffic and sends a small share to the candidate (the "challenger"). If the challenger wins, it replaces the champion alias. Inference tables record each endpoint's requests and responses, which gives you the data to compare the two.
Model Serving can serve three model types: custom models, Foundation Model APIs provisioned throughput models, and external models. Any of them can split traffic across several entries, but every served model on one endpoint must be the same type. External models have two extra rules: they must all have the same task type (for example llm/v1/chat), and each needs a unique name. Within those rules you can serve two different models, or two versions of the same model so you can try a new version while the current one stays in production.
| Served models on the endpoint | Allowed? | Example from the docs |
|---|---|---|
| Two custom models in Unity Catalog | Yes | catalog.schema.model-A v1 (current) and catalog.schema.model-B v1 (challenger) |
| Two Foundation Model APIs provisioned throughput models | Yes | meta_llama_v3_1_70b_instruct and meta_llama_v3_1_8b_instruct |
| Two external models with the same task type and unique names | Yes | gpt-4 (openai) and claude-3-opus-20240229 (anthropic), both llm/v1/chat |
| A custom model plus an external model | No | Different model types cannot share an endpoint |
Checkpoint 1 of 8· Check yourself
A data scientist wants to compare version 4 of a registered fraud model with the new version 5 on live traffic, while version 4 keeps serving production. What is the documented approach?
One endpoint can serve several versions of the same model at once, with a traffic split between them. That lets you test the new version while the current one stays in production.
“serve different versions of a model at the same time, which makes experimenting with new versions easier”Source: docs.databricks.com
Checkpoint 2 of 8· Exam question
A team wants to canary-release a new fraud model version on their live Model Serving endpoint, sending only 20% of production requests to it while the incumbent version keeps handling 80%. Both versions are already registered in Unity Catalog. How should they configure the single endpoint to achieve this split?
Correct answer: A — Add both model versions as separate served entities in the endpoint config, then set traffic_config routes so their percentages sum to 100, giving the incumbent 80 and the new version 20.
- A. Serving multiple models from one endpoint requires each version to be declared as its own served entity, and a traffic_config with routes whose percentages sum to 100 across those entities is exactly how Databricks documents A/B splits like this.
- B. Two separate endpoints have separate URLs, health checks, and scaling, and pushing the split logic into client-side random URL selection abandons the endpoint's built-in traffic routing and loses centralized control and observability of the split.
- C. A served entity maps to one model version's deployment unit; it cannot host two model artifacts internally, so there is no automatic interleaving between versions inside a single served entity.
- D. Traffic percentages are not inferred from alias history or recency; they must be explicitly set in the traffic_config routes for each served entity, so aliasing alone would not produce a controlled 80/20 split.
2.Anatomy of a split: served_entities and traffic_config
An endpoint is a REST API that exposes one or more served models. Each served model is a served entity: a named unit inside the endpoint that combines a specific model version with its own compute settings, and that can receive routed traffic. A split endpoint's configuration therefore has two parts. served_entities lists what is deployed, and traffic_config says how requests are divided among those entries.
Here is the served_entities list from the documentation's two-model custom-model example. The entity named current runs version 1 of model-A and challenger runs version 1 of model-B, each with its own workload size and scale-to-zero setting.
"served_entities":
[
{
"name":"current",
"entity_name":"catalog.schema.model-A",
"entity_version":"1",
"workload_size":"Small",
"scale_to_zero_enabled":true
},
{
"name":"challenger",
"entity_name":"catalog.schema.model-B",
"entity_version":"1",
"workload_size":"Small",
"scale_to_zero_enabled":true
}
],The split itself is in traffic_config.routes. Each route names a served entity and gives it a traffic_percentage. Watch the key name: the route refers to the entity through served_model_name, and its value must equal the name you gave the entity. It is not the Unity Catalog entity_name. In this example current receives 90% of requests and challenger receives 10%.
"traffic_config":
{
"routes":
[
{
"served_model_name":"current",
"traffic_percentage":"90"
},
{
"served_model_name":"challenger",
"traffic_percentage":"10"
}
]
}| Field | Where it sits | What it controls |
|---|---|---|
| name | Each entry in served_entities | The served entity's name inside the endpoint. Routes and direct queries refer to this name. |
| entity_name | Each entry in served_entities | The full Unity Catalog model name, e.g. catalog.schema.model-A |
| entity_version | Each entry in served_entities | Which registered version of that model this entity serves |
| workload_size / scale_to_zero_enabled | Each entry in served_entities | Compute for this entity only. Provisioned throughput entities use min_provisioned_throughput / max_provisioned_throughput instead. |
| served_model_name | Each entry in the routes list of traffic_config | Which served entity the route points to (must match its name) |
| traffic_percentage | Each entry in the routes list of traffic_config | The share of endpoint traffic that entity receives |
With a single served model there is nothing to divide. The glossary states the rule directly: traffic configuration is required when an endpoint has more than one served model. In every documented example, the route percentages add up to the endpoint's full traffic: 90/10, 60/40 and 50/50.
Checkpoint 3 of 8· Match them up
Match each Model Serving term to its meaning
Tap a term, then the definition that fits it.
The endpoint is the API, served entities are the deployed models behind it, and the traffic configuration's routes use served_model_name to give each entity its share.
“Named deployment unit inside an endpoint that represents a specific model with its compute configuration that can receive routed traffic.”Source: docs.databricks.com
3.Setting the initial split: REST API, MLflow SDK or UI
You can set the split when you create the endpoint. With the REST API, send the whole configuration, including served_entities and traffic_config, to POST /api/2.0/serving-endpoints. That is the request the 90/10 example above comes from.
With the MLflow Deployments SDK, you get a client for the Databricks target and call create_endpoint. The config dictionary uses the same keys as the REST body. The documentation's external-model example creates mix-chat-endpoint and splits traffic evenly between two served entities. Note that the Python routes give traffic_percentage as an integer, while the REST examples use quoted strings.
"traffic_config": {
"routes": [
{"served_model_name": "served_model_name_1", "traffic_percentage": 50},
{"served_model_name": "served_model_name_2", "traffic_percentage": 50}
]
},Checkpoint 4 of 8· Fill the gap
Which function returns the MLflow Deployments client used to create a split endpoint?
import mlflow.deployments
client = mlflow.deployments. ? ("databricks")
client.create_endpoint(
name="mix-chat-endpoint",mlflow.deployments.get_deploy_client("databricks") returns the client whose create_endpoint method takes the served_entities and traffic_config dictionary.
Source: docs.databricks.comThe Serving UI follows the same model. Under Serving > Create serving endpoint, the Served entities section asks for each entity's model, version and compute, and also the percentage of traffic it should receive. To add the challenger, click Add served entity and repeat those steps. Each entity has its own compute settings, so a low-traffic challenger can be sized differently from the champion. Keep in mind that scale to zero is not recommended for production endpoints, because capacity is not guaranteed after scaling to zero and the first requests have extra cold-start latency.
4.Shifting the split on a live endpoint
None of those. You update the existing endpoint's configuration with PUT /api/2.0/serving-endpoints/{name}/config. The request body has the same shape as before: the full served_entities list plus a new traffic_config. In the documentation's update example, the entities stay the same and only the routes change to 50/50. You can make the same change in the UI from the Serving tab using the Edit configuration button.
"traffic_config":
{
"routes":
[
{
"served_model_name":"current",
"traffic_percentage":"50"
},
{
"served_model_name":"challenger",
"traffic_percentage":"50"
}
]
}Clients keep calling the same endpoint while this happens. Model Serving applies configuration changes with zero downtime: the existing configuration keeps running until the new one is ready. That makes gradual rollout practical. You can move the challenger from 10% to 50%, and later make it the only served model, without a service gap or any change to calling code. Meanwhile, the endpoint's inference tables record requests and responses for each stage of the comparison.
Checkpoint 5 of 8· Check yourself
An endpoint currently routes 90% to current and 10% to challenger. Which action changes this to 50/50 as documented?
Changing the split is a configuration update on the existing endpoint, sent with PUT to the /config path. The UI equivalent is Edit configuration on the Serving tab.
“You can also make this update from the Serving tab in the Databricks UI using the Edit configuration button.”Source: docs.databricks.com
Checkpoint 6 of 8· Exam question
An endpoint is already live serving a champion and challenger model at a 50/50 traffic split. After two weeks, the challenger's business metrics look strong, so the team wants to shift the split to 90% challenger and 10% champion without any client-side changes or downtime. What is the correct way to do this?
Correct answer: A — Update the existing endpoint's traffic_config routes to the new 90/10 percentages using the update-config API or the edit configuration screen, keeping the same served entities and URL.
- A. Traffic percentages on a live endpoint are mutable through the update-config API call or the UI's edit configuration option, letting the team reweight routes between already-deployed served entities without recreating the endpoint or changing its URL.
- B. Recreating the endpoint would change its lifecycle and briefly interrupt availability, and it is unnecessary work when the same result is achieved by updating the existing traffic_config in place.
- C. Compute size affects how many requests a served entity can handle concurrently, not what share of total traffic Databricks routes to it, so resizing does not change the configured split percentage.
- D. Pushing the split logic into the client bypasses the endpoint's own traffic_config, meaning the endpoint's built-in routing, logging, and percentage enforcement no longer reflect what is actually happening.
5.Querying one served model directly
Routing by percentage is the right default for live users. Sometimes, though, you want to reach one particular model: to smoke-test the challenger before it gets any real share, or to reproduce a response from a specific model. Each served model has its own invocation path under the endpoint.
POST /serving-endpoints/{endpoint-name}/served-models/{served-model-name}/invocationsThe {served-model-name} in the path is the served entity's name (challenger), the same value the routes use as served_model_name. The request body is the same as for the endpoint, so a client can switch from split traffic to one pinned model by changing only the URL. The cost is that requests sent this way are not part of the A/B split. Use the direct path for testing and debugging, and the endpoint path for traffic that should count toward the comparison.
Checkpoint 7 of 8· Check yourself
Which statement about POST /serving-endpoints/{endpoint-name}/served-models/{served-model-name}/invocations is correct?
The request format is the same as for the endpoint, but the traffic percentages do not apply. The named served model handles the request.
“While querying the individual served model, the traffic settings are ignored.”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
Before adding a newly trained model version into an endpoint's 80/20 traffic split, an ML engineer wants to send a handful of manual test requests to only that new version to confirm it responds correctly, without affecting the live production split at all. How can this be done?
Correct answer: A — Invoke the served entity directly at its per-model invocations path, `/serving-endpoints/{endpoint-name}/served-models/{served-model-name}/invocations`, bypassing traffic_config routing.
- A. Databricks exposes a per-served-model invocations route that targets one served entity directly and skips the endpoint's traffic_config routing, which is the documented way to isolate testing to a specific model without touching the live split.
- B. Changing the traffic_config to 100% temporarily does affect the live production split during that window, since every request to the shared endpoint URL would be routed to the new version instead of the intended production mix.
- C. Standing up and tearing down a separate endpoint works but adds unnecessary provisioning overhead and is not how Databricks documents isolated testing when the model is already registered on the target endpoint.
- D. Scoring offline with `.predict` tests the model artifact in isolation but does not exercise the actual serving endpoint's REST path, authentication, or serving container, so it does not confirm the deployed served entity itself responds correctly.
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.To A/B test two real-time models you deploy them to two separate endpoints and divide traffic between the endpoints.Why is that wrong?
The documented mechanism is one endpoint that serves several models, with a traffic configuration that splits that endpoint's requests among its served entities.
Covered in Why split real-time traffic across models
2.Any two models can share an endpoint and split its traffic, for example a custom model and an external LLM.Why is that wrong?
All served models on an endpoint must be the same type: all custom models, all provisioned throughput models, or all external models with the same task type.
Covered in Why split real-time traffic across models
3.A route identifies its target by the Unity Catalog entity_name, such as catalog.schema.model-B.Why is that wrong?
A route's served_model_name must match the served entity's name inside the endpoint (for example challenger), not the registered model's full name.
Covered in Anatomy of a split: served_entities and traffic_config
4.Requests sent to a served model's /served-models/{name}/invocations path are still divided according to the endpoint's traffic percentages.Why is that wrong?
The direct path skips the split. Every request goes to the named served model.
Covered in Querying one served model directly
Practise it for real
Run a champion/challenger split on one Model Serving endpoint: create it at 90/10, move it to 50/50, then query the challenger directly.
1.POST to /api/2.0/serving-endpoints with an endpoint named multi-model whose served_entities are current (catalog.schema.model-A version 1) and challenger (catalog.schema.model-B version 1), and whose traffic_config.routes give current 90 and challenger 10.
Why: An endpoint with more than one served model needs a traffic configuration, and you can set the split when you create the endpoint.
You should see: An endpoint with two served entities that sends about 90% of requests to current and 10% to challenger.
2.PUT the same served_entities to /api/2.0/serving-endpoints/multi-model/config, this time with routes giving current 50 and challenger 50.
Why: Changing the split is a configuration update on the existing endpoint, not a new deployment.
You should see: The endpoint keeps serving on the old configuration until the new one is ready, then splits traffic evenly.
3.Send a request to /serving-endpoints/multi-model/served-models/challenger/invocations, using the same request body you would send to the endpoint.
Why: A served model's own path skips the traffic split, which is useful for testing one model on its own.
You should see: The challenger served model handles the request whatever the route percentages are.
Stuck? Get a nudge
If the routes are rejected, check that each served_model_name exactly matches the name of an entry in served_entities, not its entity_name.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/model-serving/serve-multiple-models-to-serving-endpointOfficial docs
“Serving multiple models from a single endpoint enables you to split traffic between different models to compare their performance and facilitate A/B testing.”
↩︎ Why split real-time traffic across models“You can not serve different model types in a single endpoint.”
↩︎ Why split real-time traffic across models“multiple external models in a serving endpoint as long as they all have the same task type”
↩︎ Why split real-time traffic across models“The served entity, current, hosts version 1 of model-A and gets 90% of the endpoint traffic”
↩︎ Anatomy of a split: served_entities and traffic_config“When you create model serving endpoints using the Model Serving API or the Model Serving UI, you can also set the initial traffic split”
↩︎ Setting the initial split: REST API, MLflow SDK or UI“You can also update the traffic split between served models.”
↩︎ Shifting the split on a live endpoint“In some scenarios, you might want to query individual models behind the endpoint.”
↩︎ Querying one served model directly“The request format is the same as querying the endpoint.”
↩︎ Querying one served model directly“While querying the individual served model, the traffic settings are ignored.”
↩︎ Querying one served model directly“Serving multiple models from a single endpoint enables you to split traffic between different models to compare their performance and facilitate A/B testing.”
↩︎ Exam trap 1“You cannot have both external models and non-external models in the same serving endpoint.”
↩︎ Exam trap 2“While querying the individual served model, the traffic settings are ignored.”
↩︎ Exam trap 4“For example you can not serve a custom model and an external model in the same endpoint.”
↩︎ Prediction“serve different versions of a model at the same time, which makes experimenting with new versions easier”
↩︎ Checkpoint“You can also make this update from the Serving tab in the Databricks UI using the Edit configuration button.”
↩︎ Checkpoint“then all requests are served by the challenger served model.”
↩︎ Prediction - 2.
“you might want to perform longer running online comparisons, such as A/B tests or a gradual rollout of the new model”
↩︎ Why split real-time traffic across models“An offline comparison evaluates both models against a held-out data set and tracks results using the MLflow Tracking server.”
↩︎ Why split real-time traffic across models“Model Serving and data profiling allow you to automatically collect and monitor inference tables that contain request and response data for an endpoint.”
↩︎ Why split real-time traffic across models“Model Serving executes a zero-downtime update by keeping the existing configuration running until the new one is ready.”
↩︎ Shifting the split on a live endpoint“You can create a single endpoint with multiple models and specify the endpoint traffic split between those models”
↩︎ Shifting the split on a live endpoint - 3.
“REST API that exposes one or more served models for inference.”
↩︎ Anatomy of a split: served_entities and traffic_config“Traffic configuration is required for endpoints with more than one served model.”
↩︎ Anatomy of a split: served_entities and traffic_config“Specification for what percentage of traffic to an endpoint should go to each model.”
↩︎ Key concept“Named deployment unit inside an endpoint that represents a specific model with its compute configuration that can receive routed traffic.”
↩︎ Exam trap 3“Named deployment unit inside an endpoint that represents a specific model with its compute configuration that can receive routed traffic.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/machine-learning/model-serving/create-manage-serving-endpointsOfficial docs
“Select the percentage of traffic to route to your served model.”
↩︎ Setting the initial split: REST API, MLflow SDK or UI“To add additional served entities to your endpoint, click Add served entity and repeat the configuration steps above.”
↩︎ Setting the initial split: REST API, MLflow SDK or UI“Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
↩︎ Setting the initial split: REST API, MLflow SDK or UI