TATECHATLAS
◎ English
Math & models / Guide

Choosing a regression metric for non-negative data in scikit-learn

Choose a regression metric for non-negative data by first deciding the prediction goal, then picking a consistent loss such as MAE, RMSE, or pinball loss, and finally comparing against a constant baseline with a D2-style score.

On this page

Start from the prediction target and decision goal, not from the metric list. For non-negative quantities such as counts or amounts, mean_absolute_error is a simple non-negative loss with best value 0.0 and linear cost interpretation, while root_mean_squared_error is also non-negative with best value 0.0 but penalizes larger misses more. If you need a conservative upper estimate, train a quantile regressor and evaluate it with a pinball-based scorer such as make_scorer(mean_pinball_loss, alpha=0.95). Compare against a constant baseline with a D2-style score, and evaluate several metrics together in cross-validation or search. If a scoring function is imposed externally, use it directly; otherwise prefer the same strictly consistent metric for training and evaluation.

Start from the prediction target, not from the metric list

Metric choice should follow the decision goal and the functional being predicted, not the list of available scores. In scikit-learn's model evaluation guidance, the first question is whether the scoring function is imposed externally, for example by a competition or a business contract. If it is imposed, use that scoring function directly. If you are free to choose, the guidance says to start from the ultimate goal and application of the prediction, then distinguish predicting from decision making.

For regression, the response is usually treated as a random variable with a distribution, so a point prediction usually targets a functional such as the mean, median, or a quantile. Once that functional is settled, the guidance recommends using a strictly consistent scoring function for that functional. A strictly consistent scoring function is aligned with measuring the distance between predictions and the true target functional using observations, and it can be used both as a training loss and as an evaluation metric. That alignment matters because it keeps training and evaluation speaking the same language.

For non-negative targets such as counts or amounts, this means you should decide whether you want a mean-like point forecast, a median-like point forecast, or a quantile such as an upper bound, and then pick a metric that is consistent with that choice. The metric list is secondary to that decision.

When the business impact is roughly linear in absolute error on non-negative quantities, start with mean_absolute_error because it is easy to explain to stakeholders and has best value 0.0. If large misses are disproportionately expensive, switch to root_mean_squared_error and state clearly that it penalizes large deviations more. For quantile goals such as a conservative upper estimate, train a quantile regressor and evaluate it with a pinball-based scorer at the chosen alpha, then compare against a constant baseline using a D2-style score so the improvement is measured against a simple rule rather than an arbitrary number.

Use MAE when absolute error on counts or amounts matters

Mean absolute error is a non-negative regression loss with best value 0.0, and it is straightforward to interpret because it averages absolute deviations. The scikit-learn documentation describes it as a non-negative floating point value where the best value is 0.0, and it supports sample weights and multioutput aggregation.

For non-negative targets, MAE is often a good first choice when the business impact is roughly linear in absolute error. If a miss of 3 units is about three times as costly as a miss of 1 unit, MAE matches that intuition more directly than a squared-loss metric. The worked example shows a small non-negative series where the average absolute error is 1.0.

One practical caveat is that MAE is linear in each absolute error, so it is less sensitive to very large misses than squared-error metrics, but it is not immune to them. Lower outlier sensitivity is not the same as ignoring extreme errors.

from sklearn.metrics import mean_absolute_error

y_true = [0, 2, 5, 9]
y_pred = [1, 2, 4, 10]

mae = mean_absolute_error(y_true, y_pred)
print(mae)
# 1.0

Use RMSE or root mean squared error when large errors are costly

Root mean squared error is also a non-negative regression loss with best value 0.0, and it was added to scikit-learn in version 1.4. Because it squares errors before averaging and then takes the square root, it emphasizes larger deviations more than MAE.

This makes RMSE a reasonable choice when large misses are disproportionately expensive, for example when an under-prediction or over-prediction of a non-negative amount has a nonlinear cost. The worked example uses the same inputs as the MAE example and shows a slightly larger value because the largest error contributes more under squaring.

A common misunderstanding is to treat RMSE as the standard deviation of errors in general. That equality needs special conditions such as zero mean error and a matching divisor, so it is safer to describe RMSE by its sensitivity profile rather than by equating it to an error standard deviation. Also, when translating relative changes, recompute them from the original denominator and keep the square root in every RMSE translation.

from sklearn.metrics import root_mean_squared_error

y_true = [0, 2, 5, 9]
y_pred = [1, 2, 4, 10]

rmse = root_mean_squared_error(y_true, y_pred)
print(rmse)
# 1.118033988749895

Consider quantile and pinball loss for non-symmetric risk on non-negative data

When the goal is not a central point forecast but a conservative upper estimate for non-negative demand, quantile regression and pinball loss are useful. The scikit-learn documentation shows mean_pinball_loss and explains that you can build a scorer with a specific alpha, for example make_scorer(mean_pinball_loss, alpha=0.95).

That scorer can be used to evaluate the generalization performance of a quantile regressor via cross-validation, and it can also be used for hyperparameter tuning if the sign is switched so that greater means better. The documentation also points to an example about prediction intervals for gradient boosting regression, where pinball loss is used to evaluate and tune quantile regression on data with non-symmetric noise and outliers.

This matters for non-negative data because the risk may be asymmetric: under-predicting demand can be more costly than over-predicting it, or vice versa. A quantile goal lets you target that asymmetry directly, and pinball loss gives a consistent way to evaluate it.

from sklearn.metrics import mean_pinball_loss, make_scorer

mean_pinball_loss_95p = make_scorer(mean_pinball_loss, alpha=0.95)

# Illustrative use with a quantile regressor:
# cross_val_score(estimator, X, y, cv=5, scoring=mean_pinball_loss_95p)

Use D2-style skill scores to compare against a constant baseline

D2 is described in the scikit-learn documentation as the fraction of deviance explained, and it is a generalization of R2 where squared error is replaced by a deviance of choice such as Tweedie, pinball, or mean absolute error. It is a form of skill score and is calculated as 1 minus the ratio of the model deviance to the null-model deviance.

The null prediction depends on the chosen deviance: for Tweedie it is the mean of y_true, for absolute error it is the median, and for pinball loss it is the alpha-quantile. A constant model that always predicts that null value gets a D2 of 0.0, the best possible score is 1.0, and the score can be negative because a model can be arbitrarily worse than the null model.

For non-negative regression, this is helpful because it frames improvement relative to a simple constant baseline instead of an arbitrary absolute number. If you choose a pinball-based D2, the baseline itself becomes the relevant quantile, which matches a quantile prediction goal.

Evaluate multiple metrics together in cross-validation and search

Scikit-learn permits evaluation of multiple metrics in GridSearchCV, RandomizedSearchCV, and cross_validate. The documentation describes three ways to specify multiple scoring metrics: as an iterable of string metric names, as a dict mapping a scorer name to a scoring function or predefined string, or as a callable that returns a dictionary of scores.

For non-negative regression, this is useful because one metric rarely tells the whole story. You may want MAE for linear interpretability, RMSE for sensitivity to large misses, and a pinball-based scorer for a quantile goal, all evaluated on the same cross-validation splits.

A practical note from the documentation is that custom scorers used with n_jobs greater than 1 are more robust when imported from another module rather than defined inline. Also pay attention to sign conventions when wrapping a loss as a scorer, because greater_is_better must match whether the metric is a loss or a score.

from sklearn.model_selection import cross_validate
from sklearn.metrics import mean_absolute_error, root_mean_squared_error, make_scorer

scoring = {
    'mae': make_scorer(mean_absolute_error),
    'rmse': make_scorer(root_mean_squared_error),
}

# cv_results = cross_validate(estimator, X, y, scoring=scoring, cv=5)

Use dummy estimators as a regression sanity check

The scikit-learn documentation describes dummy estimators as a simple sanity check when doing supervised learning: compare your estimator against simple rules of thumb. DummyClassifier implements several strategies for classification, and the same idea applies in spirit to regression baselines, where a constant or simple rule helps judge whether a model improves over trivial predictions.

The documentation emphasizes that with these dummy strategies the predict method completely ignores the input data, which is exactly why they are useful as baselines: they show what performance looks like when the model uses no feature information.

For non-negative regression, a constant baseline can be especially informative when paired with a D2-style score, because the null model is defined relative to the chosen deviance. If your model cannot beat a sensible constant rule, that is a strong signal to revisit the features, the target transformation, or the evaluation metric.

Align training loss, evaluation metric, and business scoring

The model evaluation guidance says that once a strictly consistent scoring function is chosen, it is best used for both training and evaluation. That recommendation is especially relevant for non-negative regression because the target functional and the cost structure should drive the metric, not convenience.

If a scoring function is imposed externally, use that one directly, even if it is not the most convenient training loss. If you are free to choose, pick a metric that matches the prediction goal and then align training and evaluation with it where possible.

In practice, this may mean using MAE or RMSE for a mean or median-oriented point forecast, using pinball loss for a quantile goal, and using a D2-style score to communicate improvement over a constant baseline. The key is that the metric should reflect the decision problem, and the same logic should guide training and evaluation whenever you have the freedom to choose.

Things to check

  • Confirm the prediction functional first: mean, median, quantile, or full distribution, because metric choice follows that choice.
  • Use mean_absolute_error when absolute error on non-negative counts or amounts matters and the business impact is roughly linear.
  • Use root_mean_squared_error when large errors are disproportionately costly, since it emphasizes larger deviations more than MAE.
  • For quantile goals, build a scorer with make_scorer(mean_pinball_loss, alpha=...) and evaluate or tune a quantile regressor with it.
  • Compare models against a constant baseline using a D2-style score, remembering that the null prediction depends on the chosen deviance.
  • Evaluate several metrics together by passing a list, dict, or callable returning a dict to GridSearchCV, RandomizedSearchCV, or cross_validate.
  • Use a dummy estimator as a regression sanity check to see whether the model improves over a trivial rule.
  • If a scoring function is imposed externally, use it directly; otherwise align training loss and evaluation metric with the same strictly consistent scoring function.

The provided sources do not give a complete decision rule for every non-negative regression case. Metric choice still depends on the application, cost structure, and the target functional being predicted. Non-negative data alone do not determine the best metric. Custom scorers and multi-metric evaluation may require extra care with parallelization and sign conventions, for example when greater_is_better must be set correctly for a loss. The examples in the sources are illustrative and may depend on dataset details, random states, or library versions.

Sources

  1. scikit-learn: model evaluation ↗
  2. scikit-learn: mean_absolute_error ↗
  3. scikit-learn: root_mean_squared_error ↗
Back to top ↑