What you will be able to do
- Explain what scikit-learn means by an estimator, and why transformers and predictors are both estimators
- Tell a transformer (fit then transform) apart from a predictor (fit then predict) by the method it exposes
- Explain why a transformer is fitted on training data only and then reused on test data
- Predict which methods a Pipeline exposes, based on its last step
- Map the same vocabulary onto Spark MLlib, where fit returns a Model that you call transform on
Key concept
The fit / transform / predict contract — Every estimator learns something from data when you call fit. A transformer then uses what it learned to turn X into a new X with transform. A predictor uses what it learned to produce outputs for new X with predict.
1.Estimator is the umbrella term: everything learns with fit
scikit-learn is the default single-node ML library on Databricks. It is included in Databricks Runtime ML, and the exam expects you to know how its API is organised. Everything in that API rests on one word: estimator. Candidates often treat "estimator" as another word for "model that predicts". That is too narrow. In scikit-learn, an estimator is any object that learns from data through a fit method. What it does after fitting decides which kind of estimator it is.
The answer is yes. A scaler has to learn the mean and standard deviation of each column before it can rescale anything, and learning from data is what makes something an estimator. There are two families you need to tell apart:
- Transformers learn a mapping from data in fit, then apply it with transform and return a new feature matrix. Examples are scalers, imputers and PCA.
- Predictors learn a relationship between X and a target y in fit, then produce outputs with predict. Examples are classifiers and regressors such as LogisticRegression or SVC. Classifiers usually also offer predict_proba for class probabilities.
The scikit-learn documentation uses this vocabulary throughout. When it describes composite models, it says transformers are combined "with other transformers or with predictors", which treats both as kinds of estimator. Many estimators also have a score method that gives a default evaluation criterion. Scoring belongs to the evaluation lessons, so we only note here that it is part of the estimator's interface too.
Checkpoint 1 of 7· Check yourself
A colleague says: "SimpleImputer isn't an estimator. It's just preprocessing." Which response is correct?
The pipeline documentation calls the non-final steps estimators and requires them to be transformers. So a transformer is a kind of estimator: it has fit, and it also has transform.
“All estimators in a pipeline, except the last one, must be transformers”Source: scikit-learn.org
Checkpoint 2 of 7· Exam question
A data scientist builds a scikit-learn `Pipeline` with steps `[('scaler', StandardScaler()), ('pca', PCA(n_components=10)), ('clf', LogisticRegression())]`. A teammate asks why `StandardScaler` and `PCA` can occupy the middle positions while `LogisticRegression` can only occupy the last position. What is the correct explanation?
Correct answer: A — Every step before the last one must implement `fit` and `transform` so its output can feed the next step, while only the final step is allowed to be a plain estimator like a classifier.
- A. This is correct: scikit-learn's `Pipeline` contract requires every step except the last to implement both `fit` and `transform`, because each step's transformed output becomes the next step's input, and only the final step is allowed to be a non-transformer such as a classifier or regressor.
- B. This is backwards; `Pipeline` does not require the last step to implement `transform` at all. When the last step is a classifier like `LogisticRegression`, the pipeline exposes `predict` instead, and calling `.transform()` on it raises an `AttributeError`.
- C. This misstates the rule: it is every step except the last that needs `transform`, not just the first, and `LogisticRegression` is deliberately exempt from implementing `transform` precisely because it is the final estimator that produces predictions.
- D. `Pipeline` executes its steps sequentially, not in parallel; each step's `fit_transform` output is passed directly as the input to the next step, which is exactly why `PCA` operates on the scaled output of `StandardScaler` rather than on raw features.
2.Transformers: fit learns the statistics, transform applies them
Raw data almost never goes straight into a model. Databricks' own feature engineering guidance notes that raw data usually needs preprocessing and transformation before a model can be built from it. In scikit-learn, that preprocessing is done by transformers, and the idea that matters is that a transformer has state.
Calling fit on a StandardScaler computes each column's mean and standard deviation and stores them on the object. Fitted attributes end in an underscore, such as mean_ and scale_. Calling transform then applies those stored numbers to whatever X you pass in. The two calls are separate on purpose: you fit once, on training data, and then transform as many datasets as you need.
>>> scaler = preprocessing.StandardScaler().fit(X_train)
>>> scaler
StandardScaler()The training set's. That is the point of keeping fit and transform apart: the test data is scaled with exactly the shifts and scales learned during fit. If you called fit again on the test set, you would get a different transformation, and your evaluation would no longer reflect how the model behaves on unseen data. Imputers follow the same contract. SimpleImputer learns a fill value per column, such as the mean, in fit, and fills gaps with those learned values in transform.
Checkpoint 3 of 7· Fill the gap
The imputer below has learned column means from its fit call. Which method fills the missing values in X with those learned means?
>>> imp = SimpleImputer(missing_values=np.nan, strategy='mean')
>>> imp.fit([[1, 2], [np.nan, 3], [7, 6]])
SimpleImputer()
>>> X = [[np.nan, 2], [6, np.nan], [7, 6]]
>>> print(imp. ? (X))SimpleImputer is a transformer. Once fitted, it is applied to data with transform. It has no predict method, because it does not model a target.
Source: scikit-learn.orgCheckpoint 4 of 7· Check yourself
You fitted a MinMaxScaler on training data. New data arrives for scoring. What is the correct way to scale it?
The fitted instance is reused, so the new data gets the same scaling and shifting that was learned on the training data.
“The same instance of the transformer can then be applied to some new test data unseen during the fit call”Source: scikit-learn.org
3.Predictors, and why a Pipeline cares which is which
A predictor also learns in fit, but it learns a relationship to a target y, and its output is a prediction, not a new feature matrix. Regressors return values from predict. Classifiers return labels from predict and usually probabilities from predict_proba.
The difference between transformers and predictors is not just naming. sklearn.pipeline.Pipeline enforces it. A pipeline is a chain of estimators, and it only works if every step except the last produces data the next step can consume. So every step except the last must be a transformer. The last step can be a predictor, a transformer, or another kind of estimator.
| Role | Examples | What fit learns | Method used after fit | Allowed position in a Pipeline |
|---|---|---|---|---|
| Transformer | StandardScaler, SimpleImputer, PCA | Parameters of a data mapping (e.g. mean_, scale_) | transform, which returns a new X | Any step |
| Predictor | LogisticRegression, SVC | A model relating X to y | predict (classifiers usually also predict_proba) | Last step only |
| Pipeline | make_pipeline(StandardScaler(), LogisticRegression()) | Fits each step in turn, feeding transformed data forward | Whatever the last step exposes | Can itself be used as an estimator |
Calling fit on a pipeline is the same as calling fit on each step in turn. Each transformer is fitted, its output is transformed and passed on, and finally the predictor is fitted on fully transformed data. At prediction time, the fitted transformers only transform. The example below chains a transformer and a predictor. One fit call learns the scaling statistics and the classifier coefficients from the training data alone.
>>> pipe = make_pipeline(StandardScaler(), LogisticRegression())
>>> pipe.fit(X_train, y_train) # apply scaling on training data
Pipeline(steps=[('standardscaler', StandardScaler()),
('logisticregression', LogisticRegression())])Checkpoint 5 of 7· Check yourself
Which of these step sequences is NOT a valid scikit-learn Pipeline?
LogisticRegression is a predictor with no transform method, so it cannot sit in a non-final position. A pipeline ending in a transformer such as PCA is valid, and it then behaves like a transformer itself.
“Pipelines require all steps except the last to be a transformer.”Source: scikit-learn.org
Checkpoint 6 of 7· Exam question
After fitting `pipe = Pipeline([('scaler', StandardScaler()), ('clf', RandomForestClassifier())])` on training data, an engineer runs `pipe.transform(X_test)` to get predictions and hits `AttributeError: 'Pipeline' object has no attribute 'transform'`. What is the most likely cause and correct fix?
Correct answer: A — The pipeline exposes the methods of its last step, and `RandomForestClassifier` is an estimator without a `transform` method, so the fix is to call `pipe.predict(X_test)` instead.
- A. This is correct: a fitted `Pipeline` only exposes the methods of its final step, and since `RandomForestClassifier` is a classifier rather than a transformer, the pipeline has no `transform` method; `predict` is the method that produces class labels from the fitted forest.
- B. This is wrong: `StandardScaler` was already fitted as part of `pipe.fit`, and even if it were unfitted the pipeline's missing `transform` is determined by the last step's type, not by an intermediate step's fit status.
- C. This is incorrect; `cross_val_score` is unrelated to which methods a `Pipeline` exposes, and running cross-validation does not add a `transform` method to a pipeline whose final estimator is a classifier.
- D. This is wrong: the `probability` parameter is specific to `SVC` and controls whether probability estimates are computed, and setting it on an unrelated classifier would not give `RandomForestClassifier`, or the pipeline wrapping it, a `transform` method.
4.The same words in Spark MLlib, with a twist
Databricks supports classic ML with both scikit-learn and Apache Spark MLlib, and the exam can use the terms "estimator" and "transformer" for either. The ideas carry over, but in the pyspark.ml Pipelines API the relationship between the two terms is arranged differently. Learn the difference explicitly:
- A Spark Transformer has a transform() method that takes a DataFrame and returns a new DataFrame, usually with extra columns appended.
- A Spark Estimator has a fit() method that takes a DataFrame and returns a Model, and that Model is itself a Transformer.
- Prediction in Spark is just another transform(). Calling transform on a fitted classification model appends a prediction column. The Pipelines API has no separate predict step.
The twist is this. In scikit-learn, a fitted StandardScaler is the same object you created, now carrying mean_ and scale_. In Spark, StandardScaler is an Estimator, and its fit() returns a separate model object that does the transforming. Likewise, a Spark Pipeline is an Estimator, and fitting it returns a PipelineModel, which is a Transformer. In both libraries, a learned step is fitted once on training data and then applied unchanged to new data. Databricks' MLlib examples include building a custom transformer, which is the extension point when the built-in feature stages are not enough.
You get back a fitted Model (a LogisticRegressionModel), which is a Transformer. You call its transform() on a DataFrame, and it returns that DataFrame with prediction columns added. Spark's Pipelines API has no predict() step in the scikit-learn sense.
Checkpoint 7 of 7· Exam question
A dataset has a categorical column `region` and a numeric column `income` that need different preprocessing before a `LogisticRegression` estimator. Which approach correctly combines heterogeneous transformers into a single scikit-learn `Pipeline` step?
Correct answer: A — Wrap `OneHotEncoder` for `region` and `StandardScaler` for `income` in a `ColumnTransformer`, then place that `ColumnTransformer` as the first step of the `Pipeline` before the classifier.
- A. This is correct: `ColumnTransformer` is designed to route different columns to different transformers, applying `OneHotEncoder` only to `region` and `StandardScaler` only to `income`, and the combined output can then be passed as the first `Pipeline` step feeding the classifier.
- B. This is wrong because applying `OneHotEncoder` to the whole DataFrame would try to one-hot-encode the numeric `income` column, and applying `StandardScaler` afterward to the whole frame would then try to scale the newly created categorical indicator columns as well.
- C. This is incorrect and impractical; scikit-learn transformers are designed to operate on full columns or arrays at once using vectorized operations, and fitting them one row at a time on single values breaks the statistics they need to learn, such as column-wide means or category sets.
- D. This is wrong: `LogisticRegression`, like other scikit-learn estimators, expects a purely numeric input matrix and has no built-in logic to detect categorical columns or encode them, so passing raw `region` strings directly would raise an error.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1."Estimator" means a model that makes predictions, so scalers and imputers are not estimators.Why is that wrong?
Estimator is the umbrella term for anything that learns with fit. Transformers are estimators that expose transform, and predictors are estimators that expose predict.
Covered in Estimator is the umbrella term: everything learns with fit
2.To scale test data correctly, call fit_transform on the test set so it is scaled by its own statistics.Why is that wrong?
A transformer is fitted on training data only. The same fitted instance then calls transform on test or new data, so both get the identical transformation.
Covered in Transformers: fit learns the statistics, transform applies them
3.Every Pipeline has a predict method.Why is that wrong?
A Pipeline has exactly the methods of its last step. If the last step is a transformer such as PCA, the pipeline behaves as a transformer and has no predict.
Covered in Predictors, and why a Pipeline cares which is which
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“scikit-learn is one of the most popular Python libraries for single-node machine learning”
↩︎ Estimator is the umbrella term: everything learns with fit - 2.https://scikit-learn.org/stable/modules/compose.htmlSecondary source
“transformers are usually combined with other transformers or with predictors (such as classifiers or regressors)”
↩︎ Estimator is the umbrella term: everything learns with fit“If the last step provides a predict method, then the pipeline would expose that method”
↩︎ Predictors, and why a Pipeline cares which is which“Calling fit on the pipeline is the same as calling fit on each estimator in turn, transform the input and pass it on”
↩︎ Predictors, and why a Pipeline cares which is which“All estimators in a pipeline, except the last one, must be transformers”
↩︎ Exam trap 1“The pipeline has all the methods that the last estimator in the pipeline has”
↩︎ Exam trap 3“All estimators in a pipeline, except the last one, must be transformers”
↩︎ Checkpoint“use all steps except the last to transform the data, and then give that transformed data to the predict method”
↩︎ Prediction“Pipelines require all steps except the last to be a transformer.”
↩︎ Checkpoint - 3.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“Estimators have a score method providing a default evaluation criterion for the problem they are designed to solve.”
↩︎ Estimator is the umbrella term: everything learns with fit“Note that for regressors, the prediction is done with predict while for classifiers it is usually predict_proba.”
↩︎ Predictors, and why a Pipeline cares which is which - 4.
“In almost all cases, the raw data requires preprocessing and transformation before it can be used to build a model.”
↩︎ Transformers: fit learns the statistics, transform applies them“ensures that the code used to compute feature values is the same during model training and when the model is used for inference”
↩︎ The same words in Spark MLlib, with a twist - 5.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“This class implements the Transformer API to compute the mean and standard deviation on a training set”
↩︎ Transformers: fit learns the statistics, transform applies them“This class implements the Transformer API to compute the mean and standard deviation on a training set”
↩︎ Key concept“The same instance of the transformer can then be applied to some new test data unseen during the fit call”
↩︎ Exam trap 2“The same instance of the transformer can then be applied to some new test data unseen during the fit call”
↩︎ Checkpoint - 6.
“Classic ML: Supervised and unsupervised learning with scikit-learn, XGBoost, LightGBM, Apache Spark MLlib, and other ML frameworks”
↩︎ The same words in Spark MLlib, with a twist - 7.
“This notebook illustrates how to create a custom transformer.”
↩︎ The same words in Spark MLlib, with a twist