TATECHATLAS
◎ English
Math & models / Guide

Evaluating Multi-output Regression in scikit-learn

Learn how to use the multioutput parameter in scikit-learn to retrieve per-target errors or apply custom weights to multi-output regression forecasts.

On this page

To evaluate a model with multiple output series in scikit-learn, use the multioutput parameter within supported metrics like mean_absolute_error or r2_score. Setting multioutput='raw_values' returns an array containing the individual score for each output stream, allowing for granular diagnosis. To calculate a single weighted score where specific outputs are more critical, pass an array-like object of weights (e.g., [0.3, 0.7]) to the multioutput parameter. This replaces the default 'uniform_average' behavior, which treats all target variables as equally important.

Understanding Multi-output Regression Metrics

In multi-target regression, a model predicts multiple continuous variables simultaneously. Standard evaluation metrics typically return a single scalar, which can hide poor performance in one specific output if others are performing well. Scikit-learn addresses this by providing mechanisms to either decompose the error per output or aggregate them using specific weighting strategies.

The choice of metric should align with the target functional. For example, if the goal is to predict the mean of a distribution, a squared loss function is appropriate. If the goal is the median, absolute loss is preferred. In multi-output scenarios, this consistency must be maintained across all target dimensions.

When using custom weights for multi-output metrics, ensure the weights sum to 1.0 to maintain the metric's original scale and interpretability as a weighted average.

Using raw_values for Per-Output Evaluation

The 'raw_values' option is essential for diagnosing which specific target variable is causing a model to underperform. Instead of averaging the results, scikit-learn returns an array where each element corresponds to the metric calculated for one output column.

This is particularly useful when target variables have different scales or different levels of noise, as it prevents the error of a high-magnitude variable from dominating the overall score.

from sklearn.metrics import mean_absolute_error; import numpy as np; y_true = np.array([[0.5, 1], [-1, 1], [7, -6]]); y_pred = np.array([[0, 2], [-1, 2], [8, -5]]); scores = mean_absolute_error(y_true, y_pred, multioutput='raw_values'); print(scores) # Illustrative output: [0.5, 1. ]

Implementing Custom Weighting for Multiple Outputs

When certain output series are more business-critical than others, a uniform average is misleading. Scikit-learn allows passing an array of weights to the multioutput parameter. The resulting score is the weighted sum of the individual errors.

For instance, if the second output is twice as important as the first, weights like [0.33, 0.67] can be applied. This ensures the final scalar reflects the prioritized importance of specific target variables.

from sklearn.metrics import mean_absolute_error; import numpy as np; y_true = np.array([[0.5, 1], [-1, 1], [7, -6]]); y_pred = np.array([[0, 2], [-1, 2], [8, -5]]); weighted_mae = mean_absolute_error(y_true, y_pred, multioutput=[0.3, 0.7]); print(weighted_mae) # Illustrative output: 0.85

Comparing Weighted vs Uniform Averaging

Uniform averaging ('uniform_average') is the default behavior, treating every output as having equal weight. This is suitable when all targets are of the same nature and importance.

In contrast, weighted averaging allows the user to penalize errors in specific dimensions more heavily. Comparing 'raw_values' against a weighted average helps determine if a global score is being skewed by a single problematic output or if the model is consistently mediocre across all targets.

Applying Multi-output Logic to Different Metrics

Not all scikit-learn metrics support the multioutput parameter. Common regression metrics such as Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and R2 score support it.

However, some robust metrics, such as median_absolute_error, do not support multi-output inputs. In such cases, users must manually iterate over the output columns and calculate the metric for each.

Creating Custom Scorers for Weighted Multi-output

To use a weighted multi-output metric within GridSearchCV or cross_val_score, you must wrap the metric using make_scorer. Since the multioutput parameter is a keyword argument, it must be passed during the scorer creation.

Because these metrics are typically losses (where lower is better), the greater_is_better parameter in make_scorer must be set to False to ensure the search algorithm correctly maximizes the negated loss.

from sklearn.metrics import make_scorer, mean_absolute_error; from sklearn.model_selection import GridSearchCV; from sklearn.linear_model import LinearRegression; weighted_mae_scorer = make_scorer(mean_absolute_error, multioutput=[0.3, 0.7], greater_is_better=False); grid = GridSearchCV(LinearRegression(), param_grid={}, scoring=weighted_mae_scorer)

Evaluating Multiple Metrics Simultaneously

For a comprehensive view, you can track multiple performance indicators using cross_validate. By passing a dictionary to the scoring parameter, you can monitor both the weighted aggregate and other metrics like R2 simultaneously.

This allows the developer to see if improving the weighted MAE comes at the cost of the overall variance explained (R2) across all outputs.

from sklearn.model_selection import cross_validate; from sklearn.metrics import make_scorer, mean_absolute_error; from sklearn.linear_model import LinearRegression; import numpy as np; X = np.random.rand(10, 3); y = np.random.rand(10, 2); scoring_dict = {'weighted_mae': make_scorer(mean_absolute_error, multioutput=[0.3, 0.7], greater_is_better=False), 'r2': 'r2'}; results = cross_validate(LinearRegression(), X, y, scoring=scoring_dict, cv=2); print(results['test_weighted_mae'])

Selecting Consistent Scoring Functions

Consistency in scoring means the metric should be aligned with the target functional of the prediction. If the model is trained to predict the mean, RMSE or MAE are appropriate.

When dealing with multi-output tasks, ensure that the weighting strategy does not inadvertently change the statistical meaning of the metric. For example, RMSE should be handled carefully when averaging, as the square root should be preserved in the translation from MSE to RMSE to maintain the original error scale.

Things to check

  • Verify that y_true and y_pred have matching shapes (n_samples, n_outputs).
  • Confirm that the chosen metric supports the multioutput parameter.
  • Ensure greater_is_better=False is set in make_scorer for error-based metrics.
  • Check that the weights array length matches the number of output columns.

The multioutput parameter is not supported by all metrics, specifically excluding median_absolute_error. It requires target and prediction arrays to have matching dimensions.

Sources

  1. scikit-learn: model evaluation ↗
  2. scikit-learn: mean_absolute_error ↗
  3. scikit-learn: root_mean_squared_error ↗
Back to top ↑