What you will be able to do
- Explain why model artifacts and the code that creates them move through environments on separate timelines
- Describe how the deploy-code pattern trains a model in development, staging and production
- Identify scenarios where promoting code is preferred, using its advantages and disadvantages
Key concept
Deploy code vs deploy models — A model is produced by code, but the model and the code can change on different schedules. So a team has to choose what it moves toward production: the training code, which then rebuilds the model in each environment, or the trained model artifact itself.
1.Code and models change on different clocks
Every model is created by code, so it is tempting to think of the two as one deliverable. In practice they drift apart. Databricks shows this with two contrasting projects. A fraud-detection pipeline retrains weekly: the code may not change very often, but the model might be retrained every week to incorporate new data. A large deep neural network for document classification is the opposite. Training it is computationally expensive and happens rarely, yet the code that deploys, serves, and monitors this model can be updated without retraining the model.
New model versions and new code don't arrive together, so an ML team has to decide which of the two it pushes through its development, staging and production environments. That choice defines two deployment patterns: deploy code (promote the training code and let each environment produce its own model) and deploy models (train once, then promote the resulting artifact).
The MLOps workflow guidance states the default plainly: during the ML development process you promote code, rather than models, from one environment to the next. It gives two reasons. First, all code in the ML process goes through the same code review and integration testing. Second, it ensures that the production version of the model is trained on production code. The Unity Catalog model-lifecycle page goes a step further. Deploying ML pipelines as code removes the need to promote models at all, because production models come from automated training workflows in the production environment.
Checkpoint 1 of 5· Check yourself
Which reason does Databricks give for promoting code, rather than models, from one environment to the next?
Promoting code means all code goes through the same review and integration testing, and the production model is trained on production code. The other options describe the opposite of what the deploy-code pattern does.
“It also ensures that the production version of the model is trained on production code.”Source: docs.databricks.com
The surrounding code. The code that deploys, serves and monitors the model can change and be updated without retraining the model. This is why the code lifecycle and the model lifecycle are treated separately.
2.How the deploy-code pattern moves through environments
In the deploy-code pattern, the code that trains models is written in the development environment, and the same code moves to staging and then to production. A model is trained in each environment, for a different purpose each time: in development as part of model development, in staging on a limited subset of data as part of integration tests, and in production on the full production data to produce the final model.
Because code is what moves, the promotion mechanism is ordinary version control. Databricks recommends storing pipelines and code in Git, so moving ML logic between stages can then be interpreted as moving code from the development branch, to the staging branch, to the release branch. MLOps Stacks builds on the same idea. Representing the model development process as code enables you to deploy the code instead of deploying the model, which makes it much easier to retrain when needed.
Checkpoint 2 of 5· Put it in order
Under the deploy-code pattern, put the places where the model is trained in the order the training code reaches them.
- 1.Production environment, on the full production data to produce the final model
- 2.Staging environment, on a limited subset of data as part of integration tests
- 3.Development environment, as part of model development
The same training code moves from development to staging to production, and each environment trains its own model: an exploratory one in development, a small-data integration test in staging, and the real model in production.
“in staging (on a limited subset of data) as part of integration tests”Source: docs.databricks.com
3.When deploy code wins, and what it costs
The deployment-patterns page lists three advantages and two disadvantages. Read them as scenario signals: if an exam question gives one of the advantages as a requirement, promoting code is the expected answer.
| Trade-off | Advantage or disadvantage | What it means in a scenario |
|---|---|---|
| Restricted access to production data | Advantage | The model is trained on production data inside the production environment, so development users never need that data |
| Automated retraining | Advantage | Safer, because the training code is reviewed, tested, and approved for production |
| Supporting code | Advantage | Follows the same pattern as training code; both go through integration tests in staging |
| Handoff of code to collaborators | Disadvantage | Steep learning curve for data scientists; predefined project templates and workflows help |
| Reviewing production training results | Disadvantage | Data scientists must be able to see results from the production environment to identify and fix ML-specific issues |
The operational-excellence best practices give more reasons why Databricks recommends deploy code for most use cases. It fits traditional software engineering workflows that use Git and CI/CD. It supports automated retraining in a locked-down environment. It requires only the production environment to have read access to production training data. It gives full control over the training environment, which makes results easier to reproduce. It also lets the data science team use modular code and iterative testing on larger projects. The same preference holds when an organisation uses several workspaces: for cross-workspace development and deployment, Databricks recommends deploying the training code to multiple environments.
Checkpoint 3 of 5· Match them up
Match each deploy-code trade-off to what the documentation says about it.
Tap a term, then the definition that fits it.
Each pair restates one item from the documented list of deploy-code advantages and disadvantages.
“Automated model retraining is safer, since the training code is reviewed, tested, and approved for production.”Source: docs.databricks.com
Checkpoint 4 of 5· Check yourself
Security policy forbids the development workspace from reading production data, and the business wants the model retrained automatically on that data. Which approach fits?
Restricted production data and safe automated retraining are both listed advantages of deploy code. The model is trained inside the production environment, which is the only environment that needs access to production data.
“In organizations where access to production data is restricted, this pattern allows the model to be trained on production data in the production environment.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A fraud detection team's compliance policy states that production customer transaction data must never leave the production Unity Catalog workspace, and no development or staging cluster may be granted read access to it. The team retrains its model weekly. Which deployment pattern satisfies this constraint?
Correct answer: A — Promote code: move the training pipeline through dev, staging, and production, and let production retrain the model against production-only data on its own schedule.
- A. This is correct because moving the training code through each environment and retraining inside production means the model is always fit on production data without that data ever being exposed to a lower environment, which is exactly what the compliance policy requires.
- B. This is incorrect because it trains on data the development workspace can access, which is not the restricted production data, so the resulting model would not reflect current production transactions and the pattern does not meet the isolation requirement.
- C. This is incorrect because bundling code and a pre-trained model together does not remove the need to train somewhere, and if that training happens outside production the data-access restriction is still violated regardless of how the artifact is packaged.
- D. This is incorrect because granting staging any access to production transaction tables, even temporary and read-only, directly violates the stated policy that production data must never leave the production workspace.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Under deploy code, the model is trained once in development and that same artifact is moved to production.Why is that wrong?
Only the code moves. A model is trained in every environment, and the production model is trained in production on the full production data.
Covered in How the deploy-code pattern moves through environments
2.Once code is promoted, production training is purely an engineering concern and data scientists no longer need to see it.Why is that wrong?
One listed disadvantage of deploy code is that data scientists must still be able to review production training results, because they are the ones who can spot ML-specific issues.
Covered in When deploy code wins, and what it costs
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The code may not change very often, but the model might be retrained every week to incorporate new data.”
↩︎ Code and models change on different clocks“the code that deploys, serves, and monitors this model can be updated without retraining the model.”
↩︎ Code and models change on different clocks“The same code moves to staging and then production.”
↩︎ How the deploy-code pattern moves through environments“The learning curve for data scientists to hand off code to collaborators can be steep.”
↩︎ When deploy code wins, and what it costs“The two patterns differ in whether the model artifact or the training code that produces the model artifact is promoted towards production.”
↩︎ Key concept“The model is trained in each environment”
↩︎ Exam trap 1“data scientists must be able to review training results from the production environment”
↩︎ Exam trap 2“In most situations, Databricks recommends the “deploy code” approach.”
↩︎ Prediction“in staging (on a limited subset of data) as part of integration tests”
↩︎ Checkpoint“Automated model retraining is safer, since the training code is reviewed, tested, and approved for production.”
↩︎ Checkpoint“In organizations where access to production data is restricted, this pattern allows the model to be trained on production data in the production environment.”
↩︎ Checkpoint - 2.
“you promote code, rather than models, from one environment to the next”
↩︎ Code and models change on different clocks“It also ensures that the production version of the model is trained on production code.”
↩︎ Code and models change on different clocks“Moving ML logic between stages can then be interpreted as moving code from the development branch, to the staging branch, to the release branch.”
↩︎ How the deploy-code pattern moves through environments - 3.
“Databricks recommends that you deploy ML pipelines as code. This eliminates the need to promote models across environments”
↩︎ Code and models change on different clocks - 4.
“Representing the model development process as code enables you to deploy the code instead of deploying the model.”
↩︎ How the deploy-code pattern moves through environments - 5.https://docs.databricks.com/aws/en/lakehouse-architecture/operational-excellence/best-practicesOfficial docs
“It requires only the production environment to have read access to prod training data.”
↩︎ When deploy code wins, and what it costs“It provides full control over the training environment, helping to simplify reproducibility.”
↩︎ When deploy code wins, and what it costs - 6.https://docs.databricks.com/aws/en/machine-learning/manage-model-lifecycle/multiple-workspacesOfficial docs
“For cross-workspace model development and deployment, Databricks recommends the deploy code approach, where the model training code is deployed to multiple environments.”
↩︎ When deploy code wins, and what it costs