CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 48/48

    Traffic Splitting on Databricks Model Serving Endpoints

    Split data between endpoints for realtime interference

    17 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why you would serve several models behind one Model Serving endpoint and split real-time traffic between them
    • Read and write the served_entities and traffic_config.routes sections of an endpoint configuration
    • Set an initial traffic split with the REST API, the MLflow Deployments SDK or the Serving UI
    • Change an existing split with PUT /api/2.0/serving-endpoints/{name}/config without recreating the endpoint
    • Send requests straight to one served model, skipping the traffic split, and know when that is useful

    Key concept

    Traffic configuration (traffic_config) — The part of a Model Serving endpoint's configuration that says what share of the endpoint's incoming requests each served model gets. You need it once an endpoint serves more than one model, and it is how real-time A/B tests and champion-versus-challenger comparisons are run.

    1.Why split real-time traffic across models

    Once a model is behind a Model Serving endpoint, every client calls one REST API. At some point you will train a better model, or think you have. There are two ways to check before switching all users to it. An offline comparison runs both models against a held-out data set and records the results with MLflow Tracking. An online comparison sends live requests to both models at once and compares how they behave in production. Databricks' MLOps guidance says real-time serving is where online comparisons, such as A/B tests or a gradual rollout, are most useful.

    Model Serving lets you run an online comparison without a second endpoint or any client changes. One endpoint can host several served models, and you set what percentage of requests each one receives. A common setup keeps the current production model (the "champion") on most of the traffic and sends a small share to the candidate (the "challenger"). If the challenger wins, it replaces the champion alias. Inference tables record each endpoint's requests and responses, which gives you the data to compare the two.

    Model Serving can serve three model types: custom models, Foundation Model APIs provisioned throughput models, and external models. Any of them can split traffic across several entries, but every served model on one endpoint must be the same type. External models have two extra rules: they must all have the same task type (for example llm/v1/chat), and each needs a unique name. Within those rules you can serve two different models, or two versions of the same model so you can try a new version while the current one stays in production.

    Which combinations can share one endpoint and split its traffic
    Served models on the endpointAllowed?Example from the docs
    Two custom models in Unity CatalogYescatalog.schema.model-A v1 (current) and catalog.schema.model-B v1 (challenger)
    Two Foundation Model APIs provisioned throughput modelsYesmeta_llama_v3_1_70b_instruct and meta_llama_v3_1_8b_instruct
    Two external models with the same task type and unique namesYesgpt-4 (openai) and claude-3-opus-20240229 (anthropic), both llm/v1/chat
    A custom model plus an external modelNoDifferent model types cannot share an endpoint

    Checkpoint 1 of 8· Check yourself

    A data scientist wants to compare version 4 of a registered fraud model with the new version 5 on live traffic, while version 4 keeps serving production. What is the documented approach?

    Checkpoint 2 of 8· Exam question

    A team wants to canary-release a new fraud model version on their live Model Serving endpoint, sending only 20% of production requests to it while the incumbent version keeps handling 80%. Both versions are already registered in Unity Catalog. How should they configure the single endpoint to achieve this split?

    Sources12

    2.Anatomy of a split: served_entities and traffic_config

    An endpoint is a REST API that exposes one or more served models. Each served model is a served entity: a named unit inside the endpoint that combines a specific model version with its own compute settings, and that can receive routed traffic. A split endpoint's configuration therefore has two parts. served_entities lists what is deployed, and traffic_config says how requests are divided among those entries.

    Here is the served_entities list from the documentation's two-model custom-model example. The entity named current runs version 1 of model-A and challenger runs version 1 of model-B, each with its own workload size and scale-to-zero setting.

    Two served entities, each pointing at a Unity Catalog model version with its own compute settingsjson
    "served_entities":
          [
             {
                "name":"current",
                "entity_name":"catalog.schema.model-A",
                "entity_version":"1",
                "workload_size":"Small",
                "scale_to_zero_enabled":true
             },
             {
                "name":"challenger",
                "entity_name":"catalog.schema.model-B",
                "entity_version":"1",
                "workload_size":"Small",
                "scale_to_zero_enabled":true
             }
          ],

    The split itself is in traffic_config.routes. Each route names a served entity and gives it a traffic_percentage. Watch the key name: the route refers to the entity through served_model_name, and its value must equal the name you gave the entity. It is not the Unity Catalog entity_name. In this example current receives 90% of requests and challenger receives 10%.

    The routes that send 90% of requests to current and 10% to challengerjson
    "traffic_config":
          {
             "routes":
             [
                {
                   "served_model_name":"current",
                   "traffic_percentage":"90"
                },
                {
                   "served_model_name":"challenger",
                   "traffic_percentage":"10"
                }
             ]
          }
    Fields involved in a traffic split, and what each one controls
    FieldWhere it sitsWhat it controls
    nameEach entry in served_entitiesThe served entity's name inside the endpoint. Routes and direct queries refer to this name.
    entity_nameEach entry in served_entitiesThe full Unity Catalog model name, e.g. catalog.schema.model-A
    entity_versionEach entry in served_entitiesWhich registered version of that model this entity serves
    workload_size / scale_to_zero_enabledEach entry in served_entitiesCompute for this entity only. Provisioned throughput entities use min_provisioned_throughput / max_provisioned_throughput instead.
    served_model_nameEach entry in the routes list of traffic_configWhich served entity the route points to (must match its name)
    traffic_percentageEach entry in the routes list of traffic_configThe share of endpoint traffic that entity receives

    With a single served model there is nothing to divide. The glossary states the rule directly: traffic configuration is required when an endpoint has more than one served model. In every documented example, the route percentages add up to the endpoint's full traffic: 90/10, 60/40 and 50/50.

    Checkpoint 3 of 8· Match them up

    Match each Model Serving term to its meaning

    Tap a term, then the definition that fits it.

    Sources31

    3.Setting the initial split: REST API, MLflow SDK or UI

    You can set the split when you create the endpoint. With the REST API, send the whole configuration, including served_entities and traffic_config, to POST /api/2.0/serving-endpoints. That is the request the 90/10 example above comes from.

    With the MLflow Deployments SDK, you get a client for the Databricks target and call create_endpoint. The config dictionary uses the same keys as the REST body. The documentation's external-model example creates mix-chat-endpoint and splits traffic evenly between two served entities. Note that the Python routes give traffic_percentage as an integer, while the REST examples use quoted strings.

    The routes from the MLflow Deployments SDK example, splitting traffic 50/50 between two external modelspython
    "traffic_config": {
                "routes": [
                    {"served_model_name": "served_model_name_1", "traffic_percentage": 50},
                    {"served_model_name": "served_model_name_2", "traffic_percentage": 50}
                ]
            },

    Checkpoint 4 of 8· Fill the gap

    Which function returns the MLflow Deployments client used to create a split endpoint?

    import mlflow.deployments
    
    client = mlflow.deployments. ? ("databricks")
    
    client.create_endpoint(
        name="mix-chat-endpoint",

    The Serving UI follows the same model. Under Serving > Create serving endpoint, the Served entities section asks for each entity's model, version and compute, and also the percentage of traffic it should receive. To add the challenger, click Add served entity and repeat those steps. Each entity has its own compute settings, so a low-traffic challenger can be sized differently from the champion. Keep in mind that scale to zero is not recommended for production endpoints, because capacity is not guaranteed after scaling to zero and the first requests have extra cold-start latency.

    Sources14

    4.Shifting the split on a live endpoint

    None of those. You update the existing endpoint's configuration with PUT /api/2.0/serving-endpoints/{name}/config. The request body has the same shape as before: the full served_entities list plus a new traffic_config. In the documentation's update example, the entities stay the same and only the routes change to 50/50. You can make the same change in the UI from the Serving tab using the Edit configuration button.

    Routes in the PUT /api/2.0/serving-endpoints/{name}/config body that move the endpoint to a 50/50 splitjson
    "traffic_config":
       {
          "routes":
          [
             {
                "served_model_name":"current",
                "traffic_percentage":"50"
             },
             {
                "served_model_name":"challenger",
                "traffic_percentage":"50"
             }
          ]
       }

    Clients keep calling the same endpoint while this happens. Model Serving applies configuration changes with zero downtime: the existing configuration keeps running until the new one is ready. That makes gradual rollout practical. You can move the challenger from 10% to 50%, and later make it the only served model, without a service gap or any change to calling code. Meanwhile, the endpoint's inference tables record requests and responses for each stage of the comparison.

    Checkpoint 5 of 8· Check yourself

    An endpoint currently routes 90% to current and 10% to challenger. Which action changes this to 50/50 as documented?

    Checkpoint 6 of 8· Exam question

    An endpoint is already live serving a champion and challenger model at a 50/50 traffic split. After two weeks, the challenger's business metrics look strong, so the team wants to shift the split to 90% challenger and 10% champion without any client-side changes or downtime. What is the correct way to do this?

    Sources12

    5.Querying one served model directly

    Routing by percentage is the right default for live users. Sometimes, though, you want to reach one particular model: to smoke-test the challenger before it gets any real share, or to reproduce a response from a specific model. Each served model has its own invocation path under the endpoint.

    The invocation path for one served model behind a multi-model endpointtext
    POST /serving-endpoints/{endpoint-name}/served-models/{served-model-name}/invocations

    The {served-model-name} in the path is the served entity's name (challenger), the same value the routes use as served_model_name. The request body is the same as for the endpoint, so a client can switch from split traffic to one pinned model by changing only the URL. The cost is that requests sent this way are not part of the A/B split. Use the direct path for testing and debugging, and the endpoint path for traffic that should count toward the comparison.

    Checkpoint 7 of 8· Check yourself

    Which statement about POST /serving-endpoints/{endpoint-name}/served-models/{served-model-name}/invocations is correct?

    Checkpoint 8 of 8· Exam question

    Before adding a newly trained model version into an endpoint's 80/20 traffic split, an ML engineer wants to send a handful of manual test requests to only that new version to confirm it responds correctly, without affecting the live production split at all. How can this be done?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.To A/B test two real-time models you deploy them to two separate endpoints and divide traffic between the endpoints.Why is that wrong?

      The documented mechanism is one endpoint that serves several models, with a traffic configuration that splits that endpoint's requests among its served entities.

      Covered in Why split real-time traffic across models

    2. 2.Any two models can share an endpoint and split its traffic, for example a custom model and an external LLM.Why is that wrong?

      All served models on an endpoint must be the same type: all custom models, all provisioned throughput models, or all external models with the same task type.

      Covered in Why split real-time traffic across models

    3. 3.A route identifies its target by the Unity Catalog entity_name, such as catalog.schema.model-B.Why is that wrong?

      A route's served_model_name must match the served entity's name inside the endpoint (for example challenger), not the registered model's full name.

      Covered in Anatomy of a split: served_entities and traffic_config

    4. 4.Requests sent to a served model's /served-models/{name}/invocations path are still divided according to the endpoint's traffic percentages.Why is that wrong?

      The direct path skips the split. Every request goes to the named served model.

      Covered in Querying one served model directly

    Practise it for real

    Run a champion/challenger split on one Model Serving endpoint: create it at 90/10, move it to 50/50, then query the challenger directly.

    1. 1.POST to /api/2.0/serving-endpoints with an endpoint named multi-model whose served_entities are current (catalog.schema.model-A version 1) and challenger (catalog.schema.model-B version 1), and whose traffic_config.routes give current 90 and challenger 10.

      Why: An endpoint with more than one served model needs a traffic configuration, and you can set the split when you create the endpoint.

      You should see: An endpoint with two served entities that sends about 90% of requests to current and 10% to challenger.

    2. 2.PUT the same served_entities to /api/2.0/serving-endpoints/multi-model/config, this time with routes giving current 50 and challenger 50.

      Why: Changing the split is a configuration update on the existing endpoint, not a new deployment.

      You should see: The endpoint keeps serving on the old configuration until the new one is ready, then splits traffic evenly.

    3. 3.Send a request to /serving-endpoints/multi-model/served-models/challenger/invocations, using the same request body you would send to the endpoint.

      Why: A served model's own path skips the traffic split, which is useful for testing one model on its own.

      You should see: The challenger served model handles the request whatever the route percentages are.

    Stuck? Get a nudge

    If the routes are rejected, check that each served_model_name exactly matches the name of an entry in served_entities, not its entity_name.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Serving multiple models from a single endpoint enables you to split traffic between different models to compare their performance and facilitate A/B testing.”
      ↩︎ Why split real-time traffic across models
      “You can not serve different model types in a single endpoint.”
      ↩︎ Why split real-time traffic across models
      “multiple external models in a serving endpoint as long as they all have the same task type”
      ↩︎ Why split real-time traffic across models
      “The served entity, current, hosts version 1 of model-A and gets 90% of the endpoint traffic”
      ↩︎ Anatomy of a split: served_entities and traffic_config
      “When you create model serving endpoints using the Model Serving API or the Model Serving UI, you can also set the initial traffic split”
      ↩︎ Setting the initial split: REST API, MLflow SDK or UI
      “You can also update the traffic split between served models.”
      ↩︎ Shifting the split on a live endpoint
      “In some scenarios, you might want to query individual models behind the endpoint.”
      ↩︎ Querying one served model directly
      “The request format is the same as querying the endpoint.”
      ↩︎ Querying one served model directly
      “While querying the individual served model, the traffic settings are ignored.”
      ↩︎ Querying one served model directly
      “Serving multiple models from a single endpoint enables you to split traffic between different models to compare their performance and facilitate A/B testing.”
      ↩︎ Exam trap 1
      “You cannot have both external models and non-external models in the same serving endpoint.”
      ↩︎ Exam trap 2
      “While querying the individual served model, the traffic settings are ignored.”
      ↩︎ Exam trap 4
      “For example you can not serve a custom model and an external model in the same endpoint.”
      ↩︎ Prediction
      “serve different versions of a model at the same time, which makes experimenting with new versions easier”
      ↩︎ Checkpoint
      “You can also make this update from the Serving tab in the Databricks UI using the Edit configuration button.”
      ↩︎ Checkpoint
      “then all requests are served by the challenger served model.”
      ↩︎ Prediction
    2. 2.
      “you might want to perform longer running online comparisons, such as A/B tests or a gradual rollout of the new model”
      ↩︎ Why split real-time traffic across models
      “An offline comparison evaluates both models against a held-out data set and tracks results using the MLflow Tracking server.”
      ↩︎ Why split real-time traffic across models
      “Model Serving and data profiling allow you to automatically collect and monitor inference tables that contain request and response data for an endpoint.”
      ↩︎ Why split real-time traffic across models
      “Model Serving executes a zero-downtime update by keeping the existing configuration running until the new one is ready.”
      ↩︎ Shifting the split on a live endpoint
      “You can create a single endpoint with multiple models and specify the endpoint traffic split between those models”
      ↩︎ Shifting the split on a live endpoint
    3. 3.
      “REST API that exposes one or more served models for inference.”
      ↩︎ Anatomy of a split: served_entities and traffic_config
      “Traffic configuration is required for endpoints with more than one served model.”
      ↩︎ Anatomy of a split: served_entities and traffic_config
      “Specification for what percentage of traffic to an endpoint should go to each model.”
      ↩︎ Key concept
      “Named deployment unit inside an endpoint that represents a specific model with its compute configuration that can receive routed traffic.”
      ↩︎ Exam trap 3
      “Named deployment unit inside an endpoint that represents a specific model with its compute configuration that can receive routed traffic.”
      ↩︎ Checkpoint
    4. 4.
      “Select the percentage of traffic to route to your served model.”
      ↩︎ Setting the initial split: REST API, MLflow SDK or UI
      “To add additional served entities to your endpoint, click Add served entity and repeat the configuration steps above.”
      ↩︎ Setting the initial split: REST API, MLflow SDK or UI
      “Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
      ↩︎ Setting the initial split: REST API, MLflow SDK or UI

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.