What you will be able to do
- Recognise the conditions under which Databricks suggests promoting a trained model artifact
- Choose between deploy code, deploy models and the hybrid approach for a given scenario
- Promote a model version across Unity Catalog environments with copy_model_version and an alias
1.Conditions that favour promoting the model artifact
In the deploy-model pattern, the model artifact is generated by training code in the development environment, then tested in staging before it is deployed to production. Databricks says to consider it when one or more of these apply: training is very expensive or hard to reproduce, all work is done in a single Databricks workspace, or the team isn't working with external repos or a CI/CD process. Its advantages follow from that list: data scientists get a simpler handoff, and when training is expensive, the model only has to be trained once.
The disadvantages mirror the strengths of deploy code. If production data can't be reached from development, which may be the case for security reasons, the pattern may not be viable at all. Automated retraining is tricky: you could automate it in development, but the production team might not accept the resulting model as production-ready. Supporting code, such as feature engineering, inference and monitoring pipelines, has to be deployed to production separately.
| Question | Deploy code | Deploy models |
|---|---|---|
| What moves toward production? | The training code | The model artifact trained in development |
| Where is the production model trained? | In production, on the full production data | In development; then tested in staging |
| Production data restricted from development? | Fits: training happens in production | May not be viable |
| Automated retraining | Safer: training code is reviewed, tested and approved | Tricky: production team may not accept the model |
| Supporting code (features, inference, monitoring) | Goes through integration tests in staging with training code | Deployed to production separately |
| Expensive training | Model is trained in each environment | Model only needs to be trained once |
Checkpoint 1 of 6· Check yourself
A small team does all its work in one Databricks workspace, has no external repo or CI/CD process, and its model takes days of GPU time to train. Which pattern does the documentation suggest considering?
Deploy code is the general default. But expensive training, a single workspace and no CI/CD are exactly the listed conditions for considering deploy models.
“You are not working with external repos or a CI/CD process.”Source: docs.databricks.com
Checkpoint 2 of 6· Exam question
An ML engineer is designing environment promotion for a new project and wants automated retraining to be safe: any change to the training logic should go through code review, unit tests, and integration tests before it is allowed to touch production. Which promotion pattern most directly delivers this guarantee?
Correct answer: A — Deploying code through each environment, since the training and supporting scripts pass through the same review, testing, and approval gates as other code before running against production data.
- A. This is correct because when training code is what moves between environments, every change to that code is reviewed, tested, and approved by the same CI/CD gates as other software before it can run in production, which is what makes automated retraining safe.
- B. This is incorrect because lineage metadata records what already happened during a past training run; it does not put the training logic itself through code review or tests before that logic is allowed to execute against production data.
- C. This is incorrect because a checksum verifies the artifact was not altered in transit, but it says nothing about whether the code that produced the artifact was reviewed or tested, so it does not deliver the requested guarantee.
- D. This is incorrect because leaving the training pipeline unmanaged means the logic that actually shapes the model's behavior skips review entirely, which is the opposite of the safety the engineer is trying to achieve.
Sources1
2.The hybrid: deploy code to staging, then deploy the model
Databricks describes a hybrid approach for exactly this requirement. If the model must be trained in staging over the full production dataset, you deploy code to staging, train the model there, and then deploy the model to production. Code is promoted for the first step and the model artifact for the second. Databricks describes the trade-off as a cost shift: the hybrid saves training costs in production but adds an extra operation cost in staging.
Checkpoint 3 of 6· Check yourself
Your requirement is that the model be trained in staging over the full production dataset. Which approach does the documentation describe for this?
This is the hybrid approach. Code goes to staging, the model is trained there on the full production data, and the model artifact is promoted to production.
“If your situation requires that the model be trained in staging over the full production dataset, you can use a hybrid approach”Source: docs.databricks.com
Sources1
3.Promoting a model version across Unity Catalog environments
When you do promote a model, Unity Catalog provides the mechanism. Typically each environment (development, staging, production) corresponds to a catalog. The model-lifecycle documentation states that deploying pipelines as code is preferred. It also acknowledges that in some cases it may be too expensive to retrain models across environments, and in those cases you copy model versions across registered models in Unity Catalog. The copy_model_version client API needs MLflow 2.8.0 or above. To run it, you need USE CATALOG on the staging and prod catalogs, USE SCHEMA on both schemas, and EXECUTE on the source model. You also need either ownership of the destination model or CREATE MODEL VERSION on it.
import mlflow
mlflow.set_registry_uri("databricks-uc")
client = mlflow.tracking.MlflowClient()
src_model_name = "staging.ml_team.fraud_detection"
src_model_version = "1"
src_model_uri = f"models:/{src_model_name}/{src_model_version}"
dst_model_name = "prod.ml_team.fraud_detection"
copied_model_version = client.copy_model_version(src_model_uri, dst_model_name)Once the copy is in the production environment, you run any pre-deployment validation and then mark the version for deployment with an alias. Governance comes from catalog permissions alone. Only users who can read the staging model and write to the production model can promote it, and the same users manage which versions are deployed. There are no separate rules or policies to configure. Unity Catalog does not support stages: the environment is expressed by the three-level name, and deployment is marked by an alias.
Checkpoint 4 of 6· Fill the gap
After validating the copied version in production, which client call marks it for deployment?
client = mlflow.tracking.MlflowClient()
client. ? (name="prod.ml_team.fraud_detection", alias="Champion", version=copied_model_version.version)Models in Unity Catalog use aliases, not stages, to mark the version for deployment. set_registered_model_alias assigns the alias to the copied version.
Source: docs.databricks.comCheckpoint 5 of 6· Put it in order
Put the steps for promoting a staging model version to production in Unity Catalog in order.
- 1.Perform any necessary pre-deployment validation in the production environment
- 2.Mark the validated version for deployment by setting an alias
- 3.Copy the staging model version into the production registered model with copy_model_version
The version has to exist in the production environment before it can be validated there, and the alias is set only after validation.
“After the model version is in the production environment, you can perform any necessary pre-deployment validation.”Source: docs.databricks.com
Checkpoint 6 of 6· Exam question
A small team runs all of its machine learning work inside a single Databricks workspace, does not use an external Git repository, and has no CI/CD pipeline configured. Training a particular model takes many hours on a large GPU cluster and is expensive to repeat. Which approach best fits this team's situation?
Correct answer: A — Promote the trained model version through Unity Catalog registered models, since one training run can be reused across environments without paying the retraining cost again or building pipeline infrastructure.
- A. This is correct because deploying the model is intended for cases where training is expensive or hard to reproduce and the team operates in a single workspace without external repos or CI/CD, so copying the already-trained model version forward avoids both the retraining cost and the infrastructure the team lacks.
- B. This is incorrect because a code-promotion structure without any pipeline to run and test that code does not give the team automated retraining benefits, and it still leaves them paying the expensive GPU training cost again for no compliance gain in this scenario.
- C. This is incorrect because retraining on a schedule ignores the stated constraint that training is expensive and hard to repeat, and repeating it in every environment multiplies that cost rather than avoiding it.
- D. This is incorrect because deploying code does not strictly require an external Git repository or CI/CD, but more importantly, building that infrastructure first does not address the team's actual constraint of expensive, hard-to-repeat training.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Deploy models is the safe choice when security keeps production data out of the development environment.Why is that wrong?
It is the reverse. Deploy models trains in development, so if production data can't be reached from there, the pattern may not work. Deploy code trains in production instead.
Covered in Conditions that favour promoting the model artifact
2.To promote a model in Unity Catalog, transition its version to the Production stage.Why is that wrong?
Unity Catalog has no stages. The catalog in the three-level name expresses the environment, and aliases promote versions for deployment.
Covered in Promoting a model version across Unity Catalog environments
Practise it for real
Promote a model version from a staging catalog to a production catalog in Unity Catalog and mark it for deployment
1.Confirm you have USE CATALOG on staging and prod, USE SCHEMA on staging.ml_team and prod.ml_team, EXECUTE on staging.ml_team.fraud_detection, and ownership or CREATE MODEL VERSION on prod.ml_team.fraud_detection.
Why: These privileges are the only governance on promotion; you don't need separate rules or policies.
You should see: The copy call has the permissions it needs.
2.Set the registry URI to databricks-uc and call client.copy_model_version(src_model_uri, dst_model_name) with MLflow 2.8.0 or above.
Why: Copying the version moves the trained artifact into the production environment without retraining.
You should see: A new model version appears under prod.ml_team.fraud_detection.
3.Run your pre-deployment validation against the copied version.
Why: Validation happens after the version is in the production environment.
You should see: You decide whether the version is fit to deploy.
4.Call client.set_registered_model_alias with alias "Champion" and the copied version number.
Why: In Unity Catalog, aliases mark which version is deployed; stages are not used.
You should see: The Champion alias on the production model now points to the copied version.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“the model artifact is generated by training code in the development environment”
↩︎ Conditions that favour promoting the model artifact“All work is done in a single Databricks workspace.”
↩︎ Conditions that favour promoting the model artifact“In cases where model training is expensive, only requires training the model once.”
↩︎ Conditions that favour promoting the model artifact“the team responsible for deploying the model in production might not accept the resulting model as production-ready.”
↩︎ Conditions that favour promoting the model artifact“Supporting code, such as pipelines used for feature engineering, inference, and monitoring, needs to be deployed to production separately.”
↩︎ Conditions that favour promoting the model artifact“deploying code to staging, training the model, and then deploying the model to production”
↩︎ The hybrid: deploy code to staging, then deploy the model“This approach saves training costs in production but adds an extra operation cost in staging.”
↩︎ The hybrid: deploy code to staging, then deploy the model“Typically an environment (development, staging, or production) corresponds to a catalog in Unity Catalog.”
↩︎ Promoting a model version across Unity Catalog environments“If production data is not accessible from the development environment (which may be true for security reasons), this architecture may not be viable.”
↩︎ Exam trap 1“Model training is very expensive or hard to reproduce.”
↩︎ Prediction“You are not working with external repos or a CI/CD process.”
↩︎ Checkpoint“If your situation requires that the model be trained in staging over the full production dataset, you can use a hybrid approach”
↩︎ Checkpoint - 2.
“in some cases, it may be too expensive to retrain models across environments.”
↩︎ Promoting a model version across Unity Catalog environments“you can copy model versions across registered models in Unity Catalog to promote them across environments.”
↩︎ Promoting a model version across Unity Catalog environments“You don't need to configure any other rules or policies to govern model promotion and deployment.”
↩︎ Promoting a model version across Unity Catalog environments“Stages are not supported for models in Unity Catalog.”
↩︎ Exam trap 2“After the model version is in the production environment, you can perform any necessary pre-deployment validation.”
↩︎ Checkpoint