What you will be able to do
- Map a scenario's objective (mean, median, quantile, probability or yes/no decision) to a scoring function that fits it
- Pick between RMSE/MSE, MAE and R² for regression, and know which of them rank models the same way
- Choose among F1, log loss, ROC AUC, precision, recall and accuracy for classification, including on imbalanced data
- Decide whether a scenario rewards recall, precision or both, and set that metric as the Databricks AutoML primary_metric
Key concept
Strictly consistent scoring function — First decide exactly what the prediction should estimate (the mean, the median, a quantile or a class probability). Then pick the scoring function that is lowest when the model gets that specific quantity right. The objective chooses the metric, not habit or the tool's default.
1.Start from the objective, not from a favourite metric
Every metric question on the exam is really a question about the objective. On Databricks, the ML lifecycle puts metric choice at the very start, in the scoping phase. Before any data is prepared, the team agrees on the prediction target, the type of problem that target implies, and how success will be measured. The lifecycle guidance asks directly: what metrics define success, whether accuracy, AUC, precision at K or business KPIs?
scikit-learn's guidance on 'Which scoring function should I use?' gives the same order of reasoning in two rules. First, if the scoring function is already set, for example by a competition or a business contract, use it. Second, if you are free to choose, start from the ultimate goal and how the prediction will be used.
That guidance then separates two activities. Predicting means estimating some property of the uncertain outcome Y given the features. Decision making means turning a prediction into an action, such as 'take an umbrella or not'. For predicting, you name the property you want, called the target functional. Typical examples are the mean, the median or a quantile. Then you pick a scoring function that is strictly consistent for it, meaning one that is optimised by predicting exactly that property. For decisions, the metrics mostly come from the confusion matrix. The next three sections apply this split to regression, then to classification probabilities, then to classification decisions.
Checkpoint 1 of 8· Check yourself
A retailer's contract with a client states that forecasts will be judged on a specific error measure. You think another metric describes model quality better. According to scikit-learn's guidance, which metric should you optimise and report?
When the scoring function is already given by the business context, the guidance is to use it. Your own choice only comes into play when you are free to choose.
“if the scoring function is given, e.g. in a kaggle competition or in a business context, use that one.”Source: scikit-learn.org
2.Regression: is the goal the mean, the median or a quantile?
scikit-learn lists, for each target functional, the scoring function that is strictly consistent for it. For regression, the question to ask is: *which summary of the outcome do I want the model to predict?*
| Functional you want to predict | Scoring or loss function | Response y | Prediction |
|---|---|---|---|
| mean | squared error | all reals | predict, all reals |
| mean | Poisson deviance | non-negative | predict, strictly positive |
| median | absolute error | all reals | predict, all reals |
| quantile | pinball loss | all reals | predict, all reals |
| mode | no consistent one exists | reals | — |
Read the table as a set of rules. If the business wants the expected value (for example, average revenue per store), squared error is the consistent choice. That covers MSE, its root RMSE, and R². A footnote to the table states that R² gives the same ranking as squared error, so in model comparison the three always pick the same winner. If the business wants the typical case, half above and half below (the median), absolute error, reported as MAE, is consistent. If the business wants a guarantee level such as 'at least 99% of days', the target is a quantile, and pinball loss is the scoring function. In the network example, the same pinball loss is used to train HistGradientBoostingRegressor(loss="quantile", quantile=0.99) and to evaluate it with mean_pinball_loss(..., alpha=0.99).
On Databricks, AutoML regression asks you to pick one of these through primary_metric. The supported values are "r2" (the default), "mae", "rmse" and "mse". A median-oriented objective therefore points to "mae". A mean-oriented objective points to the squared-error family, and within that family the three values do not change which model ranks first.
Checkpoint 2 of 8· Check yourself
A logistics team wants a delivery-time model that predicts the typical case: half of deliveries faster than the prediction, half slower. Which Databricks AutoML regression primary_metric matches this objective?
'Half faster, half slower' describes the median, and absolute error (MAE) is the strictly consistent scoring function for the median. r2, rmse and mse are all based on squared error, which measures success at predicting the mean.
“median | absolute error | all reals | predict, all reals”Source: scikit-learn.org
Checkpoint 3 of 8· Exam question
A data scientist on Databricks trains a fraud detection classifier where fraudulent transactions make up only 0.3% of the dataset. The business wants a single number that balances catching fraud against not overwhelming the review team with false alarms, and accuracy is known to be misleading on this data. Which metric should the team log with `mlflow.log_metric` as the primary evaluation score?
Correct answer: A — F1 score, computed with `sklearn.metrics.f1_score`, since it combines precision and recall into one number that stays informative when the positive class is rare
- A. F1 is the harmonic mean of precision and recall, so it stays low unless both catching fraud and limiting false alarms are handled well, which is exactly the tradeoff the business described. With a 0.3% positive rate this single number remains informative where accuracy would not.
- B. Accuracy is dominated by the 99.7% majority class here, so a model that predicts every transaction as legitimate would still score above 99% while catching zero fraud. That is the exact failure mode the business flagged as misleading.
- C. R-squared measures the proportion of variance explained in a continuous target and is a regression metric, not a classification metric, so it is not defined for a binary fraud label.
- D. Mean absolute error measures the average magnitude of numeric prediction errors and is built for regression targets, not for scoring how well a classifier separates two discrete classes.
3.Classification: score the probabilities or score the decisions?
For classifiers, the predicting-versus-deciding split becomes concrete. scikit-learn notes that for classifiers the prediction is usually predict_proba, while predict returns the decision made from those probabilities. The classification rows of the same table show which scoring functions belong to which output.
| Functional | Scoring or loss function | Prediction it scores |
|---|---|---|
| mean | Brier score | predict_proba |
| mean | log loss | predict_proba |
| mode | zero-one loss | predict, categorical |
When the probabilities themselves matter, for example a risk score that a downstream system reads directly, use log loss. Log loss is also called logistic regression loss or cross-entropy loss. It is defined on probability estimates and evaluates predict_proba outputs rather than discrete labels.
When you need to rank or separate the classes without choosing a cut-off yet, use ROC AUC. roc_auc_score summarises the ROC curve in one number, and unlike the F1 score, ROC does not require you to optimise a threshold.
When the output is a hard decision, most of the relevant scores are built from the confusion matrix. Zero-one loss is one minus accuracy, so the two give different values but the same ranking. Accuracy has a weakness, though. scikit-learn provides balanced_accuracy_score specifically because plain accuracy can give inflated performance estimates on imbalanced datasets. Balanced accuracy is the average of per-class recall, and on a balanced dataset it equals ordinary accuracy. When one class is rare and you care about the positive class, the precision, recall and F-measure family applies. scikit-learn describes F1 as the 'balanced F-score'.
Databricks AutoML classification turns this into one parameter. primary_metric accepts "f1" (the default), "log_loss", "precision", "accuracy" and "roc_auc". In binary problems, pos_label tells AutoML which class is positive, which precision and recall need in order to be calculated.
Checkpoint 4 of 8· Fill the gap
This is the start of the AutoML classify signature. Which value is the default for primary_metric?
databricks.automl.classify(
dataset: Union[pyspark.sql.DataFrame, pandas.DataFrame, pyspark.pandas.DataFrame, str],
*,
target_col: str,
primary_metric: str = " ? ",AutoML classification ranks runs by F1 unless you set something else. The default is not necessarily right for your scenario. If the objective is well-calibrated probabilities, log_loss is the metric that fits.
Source: docs.databricks.comCheckpoint 5 of 8· Exam question
A hospital is training a model on Databricks to flag patients who may have a rare but life-threatening condition, feeding a follow-up diagnostic test. Missing a true case is far more costly than sending a healthy patient for an unnecessary follow-up test. Which metric should the team optimize as the primary decision criterion?
Correct answer: A — Recall, computed with `sklearn.metrics.recall_score`, since it measures the share of true positives the model successfully identifies out of all actual positives
- A. Recall is true positives divided by all actual positives, so it directly measures how many real cases of the condition the model catches. Optimizing recall minimizes missed diagnoses, which matches a scenario where false negatives are the costly error.
- B. Precision measures how trustworthy a positive prediction is, which matters when false positives are expensive, but here an unnecessary follow-up test is described as the cheaper error, so precision is not the priority metric.
- C. Root mean squared error penalizes numeric prediction error and applies to continuous regression targets, not to a binary diagnostic classification decision.
- D. Specificity measures how well the model avoids false alarms on healthy patients, which is the opposite priority from the scenario, where minimizing missed true cases matters far more than minimizing unnecessary follow-ups.
4.Precision or recall: what does a miss cost compared with a false alarm?
Databricks' evaluation guidance defines the two terms in plain words. Precision asks: of the items I returned, what percentage are actually relevant? Recall asks: of all the items I know are relevant, what percentage did I return? In the worked example, two of three returned results are relevant, so precision is 0.66 (2/3). The results cover two of four relevant documents, so recall is 0.5 (2/4).
The definitions also differ in what they need from your labels. Precision can be computed without knowing every relevant item. Recall needs ground truth that contains all the relevant items.
Which one to prioritise depends on the scenario. Databricks' retrieval-quality guide sorts use cases by this question, and the same reasoning applies to any classifier with a positive class:
- Recall first, when missing something is unacceptable. Examples: clinical-trial matching that cannot miss eligible patients, compliance search that needs every relevant regulation, root-cause analysis that must surface all related incidents. - Precision first, when only the most relevant results should come through. Examples: entity resolution, flagging duplicate transactions with high confidence, support engineers who need the exact solution at the top. - Both, when a case needs coverage and accuracy at once. Example: M&A due diligence, which cannot miss risks but also needs relevant documents first. For a classifier, this is where F1, the balanced F-score, fits.
Checkpoint 6 of 8· Match them up
Match each scenario to the metric emphasis the Databricks guidance assigns it
Tap a term, then the definition that fits it.
The deciding question is which error the business can least afford. Missing an eligible patient is a recall failure. A wrong duplicate flag is a precision failure. Due diligence cannot afford either.
“Pharma clinical trial matching: Cannot miss eligible patients or relevant studies.”Source: docs.databricks.com
Checkpoint 7 of 8· Exam question
A marketing team is evaluating several candidate classifiers trained in a Databricks notebook and has not yet decided on a probability cutoff for ranking leads, since the sales team wants the flexibility to adjust the cutoff later based on capacity. Which metric best compares the classifiers' ability to rank positive cases above negative cases across all possible thresholds?
Correct answer: A — ROC AUC, computed with `sklearn.metrics.roc_auc_score`, since it summarizes ranking quality across every possible classification threshold in one number
- A. ROC AUC evaluates the true positive rate against the false positive rate across every threshold, so it measures how well a model ranks positives above negatives without committing to a specific cutoff. That matches a scenario where the cutoff is deliberately left open.
- B. F1 at a fixed 0.5 threshold only reflects performance at that one cutoff, so it does not answer how the models compare once the sales team later shifts the cutoff based on capacity.
- C. Accuracy at a single fixed threshold has the same limitation as F1 at that threshold: it reflects one operating point and hides how ranking quality changes as the cutoff moves.
- D. Mean absolute error against binary 0/1 labels is not a standard way to evaluate classifier ranking quality and does not summarize threshold-independent separation between classes the way a ranking metric does.
5.Once chosen, use the metric everywhere
Choosing the metric is not a reporting detail added at the end. scikit-learn advises that once a strictly consistent scoring function is chosen, it is best used both as the training loss and as the metric for evaluating and comparing models. The network example does exactly that: pinball loss for training, for hyperparameter search and for comparing models. In scikit-learn, cross-validation tools such as GridSearchCV take the metric through their scoring parameter. If you skip that, an estimator's score method falls back to a default criterion, which is accuracy for most classifiers and R² for most regressors, whatever your objective is. Dummy estimators give a baseline value of your chosen metric for random predictions, so you can tell whether a score is actually good.
In Databricks AutoML, the metric you select is the primary metric used to score and rank every run. AutoML also stops training and tuning early when that validation metric stops improving. A badly chosen metric therefore changes which model wins *and* when the search stops. Finally, the lifecycle guidance recommends logging metrics to MLflow runs, and notes that the metrics defined during development can be reused for production monitoring. The objective you scoped on day one then remains the yardstick in production.
Checkpoint 8 of 8· Check yourself
You have chosen pinball loss at alpha=0.99 because the objective is a 99% quantile. Where does scikit-learn's guidance say this scoring function is best used?
Using the same strictly consistent function for training, tuning and comparison keeps every stage optimising the objective the business actually set.
“it is best used for both: as loss function for model training and as metric/score in model evaluation and model comparison.”Source: scikit-learn.org
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The tool's default metric (for example AutoML's f1 or r2) is a safe choice for any scenario.Why is that wrong?
Defaults are only starting points. When you are free to choose, the metric should follow from the goal and how the prediction will be used.
Covered in Start from the objective, not from a favourite metric
2.Switching from RMSE to R² can change which regression model ranks best.Why is that wrong?
R² ranks models in the same order as squared error. Changing between them changes the numbers, not the winner. To change what is being optimised you need a different functional, such as the median (MAE) or a quantile (pinball loss).
Covered in Regression: is the goal the mean, the median or a quantile?
3.Plain accuracy is a reliable headline metric even when one class is rare.Why is that wrong?
On imbalanced data, accuracy can be inflated. scikit-learn provides balanced accuracy to avoid this, and the precision/recall/F1 family focuses on the positive class.
Covered in Classification: score the probabilities or score the decisions?
4.Recall, like precision, can be computed from the returned results alone.Why is that wrong?
Precision only needs the relevance of what was returned. Recall needs ground truth that lists every relevant item.
Covered in Precision or recall: what does a miss cost compared with a false alarm?
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“What metrics define success: accuracy, AUC, precision at K, or business KPIs?”
↩︎ Start from the objective, not from a favourite metric“Before building anything, align on what the model needs to do and how you will know it is working.”
↩︎ Start from the objective, not from a favourite metric“The metrics you define during development and training can be reused later as metrics for production monitoring.”
↩︎ Once chosen, use the metric everywhere - 2.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“Typical examples are the mean (expected value), the median or a quantile of the response variable”
↩︎ Start from the objective, not from a favourite metric“most of them are covered with or derived from the metrics.confusion_matrix”
↩︎ Start from the objective, not from a favourite metric“R² gives the same ranking as squared error.”
↩︎ Regression: is the goal the mean, the median or a quantile?“Log loss, also called logistic regression loss or cross-entropy loss, is defined on probability estimates.”
↩︎ Classification: score the probabilities or score the decisions?“ROC doesn’t require optimizing a threshold for each label.”
↩︎ Classification: score the probabilities or score the decisions?“The zero-one loss is equivalent to one minus the accuracy score, meaning it gives different score values but the same ranking.”
↩︎ Classification: score the probabilities or score the decisions?“Compute the F1 score, also known as balanced F-score or F-measure.”
↩︎ Classification: score the probabilities or score the decisions?“Dummy estimators are useful to get a baseline value of those metrics for random predictions.”
↩︎ Once chosen, use the metric everywhere“Estimators have a score method providing a default evaluation criterion for the problem they are designed to solve.”
↩︎ Once chosen, use the metric everywhere“use a strictly consistent scoring function for that (target) functional”
↩︎ Key concept“If you are free to choose, it starts by considering the ultimate goal and application of the prediction.”
↩︎ Exam trap 1“R² gives the same ranking as squared error.”
↩︎ Exam trap 2“The balanced_accuracy_score function computes the balanced accuracy, which avoids inflated performance estimates on imbalanced datasets.”
↩︎ Exam trap 3“if the scoring function is given, e.g. in a kaggle competition or in a business context, use that one.”
↩︎ Checkpoint“So the target functional is the 99% quantile. From the table above, you choose the pinball loss as scoring function”
↩︎ Prediction“median | absolute error | all reals | predict, all reals”
↩︎ Checkpoint“it is best used for both: as loss function for model training and as metric/score in model evaluation and model comparison.”
↩︎ Checkpoint - 3.
“Supported metrics for regression: “r2” (default), “mae”, “rmse”, “mse””
↩︎ Regression: is the goal the mean, the median or a quantile?“Supported metrics for classification: “f1” (default), “log_loss”, “precision”, “accuracy”, “roc_auc””
↩︎ Classification: score the probabilities or score the decisions?“(Classification only) The positive class. This is useful for calculating metrics such as precision and recall.”
↩︎ Classification: score the probabilities or score the decisions? - 4.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“two out of the three retrieved results were relevant to the user's query, so the precision was 0.66 (2/3).”
↩︎ Precision or recall: what does a miss cost compared with a false alarm?“Computing precision does not require knowing all relevant items.”
↩︎ Precision or recall: what does a miss cost compared with a false alarm?“Computing recall requires your ground-truth to contain all relevant items.”
↩︎ Exam trap 4 - 5.
“If precision matters most (need only the most relevant results):”
↩︎ Precision or recall: what does a miss cost compared with a false alarm?“M&A due diligence: Can't miss risks (recall) but need relevant docs first (precision).”
↩︎ Precision or recall: what does a miss cost compared with a false alarm?“Pharma clinical trial matching: Cannot miss eligible patients or relevant studies.”
↩︎ Checkpoint - 6.
“The evaluation metric is the primary metric used to score the runs.”
↩︎ Once chosen, use the metric everywhere“it stops training and tuning models if the validation metric is no longer improving.”
↩︎ Once chosen, use the metric everywhere