What you will be able to do
- Decide whether a model from another platform can be logged as a built-in type or has to be wrapped in CustomModel
- Log an MLflow pyfunc model and control which MLflow metadata and dependencies come with it
- Wrap a pickled or file-based model in a CustomModel with ModelContext and code_paths
- Pick dependency and target-platform settings so a migrated model runs where you need it
- Map the jobs an external MLOps stack does (experiment tracking, features, datasets, monitoring, orchestration, lineage) to the Snowflake-native capability that covers each
Key concept
CustomModel as the migration bridge — If a model from another platform is not one of the registry's built-in types, you wrap it in a CustomModel subclass. That subclass loads the model's files or objects from a ModelContext and exposes inference methods, and the wrapper gets logged like any other model.
1.First decision: built-in type or custom wrapper?
Any model you move into Snowflake ends up in the Snowflake Model Registry. You log the model, and the registry stores it as a first-class schema-level object with versions, metadata and role-based access. The registry is built to accept models wherever they came from: it is "flexible and powerful enough to support your own previously-trained models" along with custom processing code. Where the model was trained matters less than what kind of Python object it is.
So the first step in any migration is to check the model's framework against the built-in list. If it is on the list, you pass the in-memory model object straight to log_model and Snowflake handles serialization. If it is not, you wrap it in a CustomModel.
| What you are migrating | Logging path |
|---|---|
| scikit-learn, XGBoost, LightGBM, CatBoost, Prophet models | Built-in type: pass the model object to log_model |
| PyTorch, TensorFlow, Keras models | Built-in type: pass the model object to log_model |
| Hugging Face pipeline, Sentence Transformer | Built-in type: pass the model object to log_model |
| MLFlow PyFunc models from an MLflow tracking server | Built-in type, with MLflow-specific log_model options |
| Any other serializable model or file-based artifact | snowflake.ml.model.CustomModel wrapper |
Snowsight offers a second way in. The registry overview notes that "You can also import a model from an external provider to Snowflake," and in the Snowsight import-and-deploy flow you name the version, choose a database and schema, and can add pip requirements. Only packages served from PyPI are supported there. The flow then continues straight into deploying a service.
Checkpoint 1 of 6· Check yourself
A model built with a framework that is not on the registry's built-in list needs to move into Snowflake. What does the registry offer for it?
The built-in models page directs every other model type to the CustomModel class, which wraps serialized models and custom code.
“Other types of models are supported via the snowflake.ml.model.CustomModel class”Source: docs.snowflake.com
2.Bringing MLflow pyfunc models across
Teams leaving an MLflow-based setup often have models already logged as MLflow pyfunc. Snowflake accepts these directly: "You can use MLflow models that support PyFunc." If the MLflow model has a signature, the registry infers it. If not, you have to supply either a signature or sample_input_data.
What needs deciding during migration is how much of MLflow's packaging comes along. Three keys in the options dictionary of log_model control this.
| Option | What it controls | Default |
|---|---|---|
| model_uri | URI of the MLflow model's artifacts; required if it is not in the model's metadata | n/a |
| ignore_mlflow_metadata | If True, the model's MLflow metadata is not imported into the registry model object | False |
| ignore_mlflow_dependencies | If True, the dependencies in the model's MLflow metadata are ignored | False |
The dependency switch matters most. An MLflow model records the packages from the environment it was trained in, and Snowflake warehouses may not offer those packages. The documented pattern is to ignore MLflow's dependency list and state your own with conda_dependencies.
Checkpoint 2 of 6· Fill the gap
Which option key makes the registry drop the dependency list recorded in the MLflow model's metadata?
model_ref = registry.log_model(
mlflow.pyfunc.load_model(f"runs:/{run_id}/model"),
model_name="mlflowModel",
version_name="v1",
conda_dependencies=["mlflow<=2.4.0", "scikit-learn", "scipy"],
options={" ? ": True}
)
model_ref.run(X_test)ignore_mlflow_dependencies discards MLflow's recorded packages so the explicit conda_dependencies list is used. ignore_mlflow_metadata covers metadata, not packages.
Source: docs.snowflake.comCheckpoint 3 of 6· Exam question
A team trained a scikit-learn model in Amazon SageMaker and has the joblib artifact. They want analysts to score Snowflake tables with SQL inside a warehouse. Which migration path fits?
Correct answer: A — Load the model with the training scikit-learn version, call `Registry.log_model` with `sample_input_data` and `conda_dependencies`, then score via `MODEL!PREDICT`.
- A. Correct: log_model captures the model, its signature (from sample_input_data) and pinned dependencies, so the registered version runs in a warehouse and is callable from SQL via MODEL!PREDICT.
- B. Incorrect: the registry is not limited to Snowflake-trained models, and a per-row UDF that deserializes a file is slow and bypasses versioning, signatures and governance.
- C. Incorrect: warehouses can run scikit-learn models registered from elsewhere; a copied container in SPCS adds infrastructure and is not required for this scenario.
- D. Incorrect: keeping SageMaker batch transform leaves the external platform in place and produces stale precomputed predictions instead of migrating the model.
Sources3
3.Wrapping pickles and other artifacts with CustomModel
For anything outside the built-in list, you write a subclass of custom_model.CustomModel. Its ModelContext holds what the model needs at inference time, either as in-memory objects of supported types or as file paths. File paths are the usual migration case, because the old platform usually hands over artifacts on disk: "For models or data that are stored in serialized files like Python pickles or JSON, you can provide file paths to your ModelContext." The files get packaged with the model. Your class loads them and exposes methods decorated with @custom_model.inference_api.
# Initialize ModelContext with a file path
# my_file_path is a local pickle file path
model_context = custom_model.ModelContext(
my_file_path='/path/to/file.pkl',
)
# Define a custom model class that loads the pickled object
class ExampleBringYourOwnModel(custom_model.CustomModel):
def __init__(self, context: custom_model.ModelContext) -> None:
super().__init__(context)
# Use 'my_file_path' key from the context to load the pickled object
with open(self.context['my_file_path'], 'rb') as f:
self.obj = pickle.load(f)
@custom_model.inference_api
def predict(self, input: pd.DataFrame) -> pd.DataFrame:
# Use the loaded object to make predictions
model_output = self.obj.predict(input)
return pd.DataFrame({'output': model_output})Three rules prevent common migration bugs. First, if the bundle includes a supported model such as XGBoost alongside unsupported pieces, put that model object directly in the context and Snowflake serializes it for you. Second, always read models through self.context[...]. Assigning an in-memory model directly to an attribute captures a second copy in a closure and makes the model files much larger. Third, write inference methods for multi-row DataFrames. Server-side batching can merge single-record requests into one DataFrame, so a predict function that assumes one row will break once the model is serving real-time traffic.
After the wrapper works locally, you log it like any other model, with explicit dependencies and sample input to infer the signature.
reg = Registry(session=session, database_name="ML", schema_name="REGISTRY")
mv = reg.log_model(my_model,
model_name="my_custom_model",
version_name="v1",
conda_dependencies=["scikit-learn"],
comment="My Custom ML Model",
sample_input_data=train_features)
output_df = mv.run(input_df)Migrated models often rely on helper modules from the old codebase, such as preprocessing or feature utilities. The code_paths parameter packages that code with the model so it can be imported the same way as locally. With string paths, the last component of each path becomes the module name. A CodePath object with a filter lets you include only part of a directory tree, for example leaving out a tests/ folder.
mv = reg.log_model(
my_model,
model_name="my_model",
version_name="v1",
code_paths=["src/mymodule"], # import with: import mymodule
)Checkpoint 4 of 6· Put it in order
Put the steps for migrating a pickled model as a CustomModel in order
- 1.Log it with reg.log_model, supplying dependencies and sample_input_data
- 2.Test the model locally by calling its predict method on a DataFrame
- 3.Instantiate the CustomModel subclass with a ModelContext pointing at the pickle file
- 4.Run inference on the logged version with mv.run
The model has to exist before you can test it. The docs say to log only after it works as intended, and mv.run needs the ModelVersion that log_model returns.
“When the model works as intended, log it to the Snowflake Model Registry.”Source: docs.snowflake.com
Sources4
4.Choosing where the migrated model runs
On the old platform the model probably ran in a container you configured yourself. In Snowflake it runs either in a virtual warehouse or in Snowpark Container Services (SPCS), and you choose between them with target_platforms. The choice changes how dependencies are handled. SPCS is also where GPU inference happens, since the registry can "serve the model in Snowpark Container Services for GPU-based inference." Warehouse-deployed models have a size limit of 15 GB.
| target_platforms | Model kind | Default behavior of log_model() |
|---|---|---|
| SNOWPARK_CONTAINER_SERVICES | Built-in model type | pip_requirements populated automatically from the environment; not runnable in WAREHOUSE |
| SNOWPARK_CONTAINER_SERVICES | Custom Model | You must provide all dependencies in conda_dependencies or pip_requirements; not runnable in WAREHOUSE |
| WAREHOUSE | Built-in model type | conda_dependencies populated automatically; log_model() fails if the model is not runnable in WAREHOUSE |
| WAREHOUSE | Custom Model | log_model() fails if the model is not runnable in WAREHOUSE; you must provide the dependencies |
For migrations, this means a custom-wrapped model gets no automatic dependency inference: you list every package it imports. If those packages come from PyPI, including a private repository your old platform used, set artifact_repository_map so pip packages install from that repository.
Checkpoint 5 of 6· Check yourself
A team logs a CustomModel wrapper with target_platforms set to SNOWPARK_CONTAINER_SERVICES only. Which statement is true?
Automatic pip_requirements population applies to built-in types. Custom models targeting SPCS must declare their dependencies and cannot run in the warehouse.
“Users must provide all dependencies in either of conda_dependencies and pip_requirements.”Source: docs.snowflake.com
Sources1
5.Snowflake-native alternatives to external MLOps tools
Migrating the model is only part of leaving an external MLOps stack. The surrounding jobs (tracking experiments, managing features, versioning data, monitoring, orchestrating pipelines, tracing lineage) each have a Snowflake-native capability, so the workflow can stay where the data is. Snowflake ML is described as flexible and modular, so you can adopt these pieces one at a time rather than all at once.
Experiments are a good example for teams coming from an external tracking tool. Snowflake ML Experiments organize training runs, each with its own metrics, parameters and artifacts, and the results are visible in Snowsight. You can log runs during training on Snowflake, or upload metadata and artifacts from training that already happened elsewhere. The Feature Store can also work with external tools: user-managed feature pipelines can use tools such as dbt.
| Job in an external MLOps stack | Snowflake-native capability |
|---|---|
| Experiment tracking and comparing runs | Snowflake ML Experiments |
| Feature engineering and a central feature repository | Snowflake Feature Store |
| Versioned, reproducible training data | Snowflake Datasets |
| Model storage, versions and metadata | Snowflake Model Registry |
| Monitoring performance and drift of production models | ML Observability |
| Computing Shapley-value explanations | ML Explainability |
| Scheduling and chaining retraining pipelines | Snowflake Task Graphs (DAGs) |
| Tracing source data to features, datasets and models | ML Lineage |
| Running training scripts from an external IDE | Snowflake ML Jobs on Container Runtime |
Checkpoint 6 of 6· Check yourself
A team used an external tool to compare hyperparameter and metric results across many training runs. Which Snowflake-native capability covers that job?
Experiments are made for comparing runs, each with its own metrics, parameters and artifacts. Datasets version data, Observability watches deployed models, and Lineage traces provenance.
“With Snowflake ML Experiments, you can set up experiments, organized evaluations of the results of model training.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A scikit-learn or XGBoost model trained on another platform has to be wrapped in CustomModel because it came from outside Snowflake.Why is that wrong?
Built-in support depends on the model's framework, not where it was trained. Log the model object directly and keep CustomModel for types that are not supported.
2.An MLflow model's recorded dependencies always carry over cleanly, so you never need to override them.Why is that wrong?
The packages MLflow recorded may not be available in Snowflake warehouses. That is why ignore_mlflow_dependencies exists, so you can supply your own list instead.
Covered in Bringing MLflow pyfunc models across
3.Inside a CustomModel it is fine to assign the in-memory model directly, as in self.model = model.Why is that wrong?
Reading the model through the context avoids capturing a second copy in a closure, which would make the serialized model files much larger.
Covered in Wrapping pickles and other artifacts with CustomModel
4.Moving to Snowflake means every external MLOps tool must be replaced, so user-managed pipelines with tools like dbt are not possible.Why is that wrong?
Snowflake ML is modular. The Feature Store explicitly supports user-managed feature pipelines with external tools such as dbt, and ML Jobs let teams keep working from an external IDE.
Covered in Snowflake-native alternatives to external MLOps tools
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The Model Registry is also flexible and powerful enough to support your own previously-trained models, as well as any custom processing code.”
↩︎ First decision: built-in type or custom wrapper?“You can also import a model from an external provider to Snowflake.”
↩︎ First decision: built-in type or custom wrapper?“serve the model in Snowpark Container Services for GPU-based inference”
↩︎ Choosing where the migrated model runs“Maximum total model size of 15 GB (for warehouse deployed models)”
↩︎ Choosing where the migrated model runs“To install pip packages from a private or built-in PyPI repository, set artifact_repository_map.”
↩︎ Choosing where the migrated model runs“The Snowflake Model Registry has built-in type support for the most common model types, including scikit-learn, xgboost, LightGBM”
↩︎ Exam trap 1“The Snowflake Model Registry has built-in type support for the most common model types, including scikit-learn, xgboost, LightGBM”
↩︎ Prediction“Users must provide all dependencies in either of conda_dependencies and pip_requirements.”
↩︎ Checkpoint - 2.
“Only packages served from PyPi are supported.”
↩︎ First decision: built-in type or custom wrapper? - 3.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/built-in-models/mlflowOfficial docs
“You can use MLflow models that support PyFunc.”
↩︎ Bringing MLflow pyfunc models across“Otherwise, you must provide either signature or sample_input_data.”
↩︎ Bringing MLflow pyfunc models across“which is useful due to package available limitations in Snowflake warehouses”
↩︎ Exam trap 2 - 4.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/bring-your-own-model-typesOfficial docs
“For models or data that are stored in serialized files like Python pickles or JSON, you can provide file paths to your ModelContext.”
↩︎ Wrapping pickles and other artifacts with CustomModel“Don’t assume that the input DataFrame will always contain a single row.”
↩︎ Wrapping pickles and other artifacts with CustomModel“The last component of each path becomes the importable module name.”
↩︎ Wrapping pickles and other artifacts with CustomModel“Serializable models trained using external tools or obtained from open source repositories can be used with CustomModel.”
↩︎ Key concept“In your custom model class, always access model objects through the model context.”
↩︎ Exam trap 3“When the model works as intended, log it to the Snowflake Model Registry.”
↩︎ Checkpoint - 5.
“With Snowflake ML Experiments, you can set up experiments, organized evaluations of the results of model training.”
↩︎ Snowflake-native alternatives to external MLOps tools - 6.
“you can upload your own metadata and artifacts from prior training”
↩︎ Snowflake-native alternatives to external MLOps tools“teams that prefer working from an external IDE (VS Code, PyCharm, SageMaker Notebooks) to dispatch functions, files or modules”
↩︎ Snowflake-native alternatives to external MLOps tools“ML Lineage is a capability to trace end-to-end lineage of ML artifacts from source data to features, datasets, and models.”
↩︎ Snowflake-native alternatives to external MLOps tools - 7.
“Ability to use user-managed feature pipelines with external tools such as dbt”
↩︎ Snowflake-native alternatives to external MLOps tools“Ability to use user-managed feature pipelines with external tools such as dbt”
↩︎ Exam trap 4 - 8.
“Each version holds a materialized snapshot of your data with guaranteed immutability”
↩︎ Snowflake-native alternatives to external MLOps tools - 9.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-observabilityOfficial docs
“ML Observability allows you to track the quality of production models you have deployed via the Snowflake Model Registry”
↩︎ Snowflake-native alternatives to external MLOps tools - 10.
“Use CoCo to build Snowflake Task Graphs (DAGs) that chain feature refresh, training, evaluation, and deployment.”
↩︎ Snowflake-native alternatives to external MLOps tools
Also cited
- https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/built-in-models/overviewOfficial docs
“Other types of models are supported via the snowflake.ml.model.CustomModel class”
↩︎ Checkpoint