Quantile Regression with Pinball Loss in scikit-learn
Learn how to implement, evaluate, and tune quantile regression models using pinball loss and the D² skill score in scikit-learn.
On this page
The short answer
To perform quantile regression in scikit-learn, you must align your model's objective with your evaluation metric by selecting the same target quantile (alpha). You can use regressors like HistGradientBoostingRegressor with loss='quantile' and specify the target via the 'quantile' parameter. For evaluation, use mean_pinball_loss with the corresponding 'alpha' parameter. To use these metrics in cross-validation or hyperparameter tuning, wrap them with make_scorer, ensuring you set greater_is_better=False because pinball loss is a value to be minimized. For assessing model skill relative to a baseline, use d2_pinball_score, which generalizes the R² concept to quantiles. When working with multiple targets, use the multioutput parameter to define how errors are aggregated across outputs.
Define the target quantile for your decision
Before training, you must identify the specific quantile (alpha) required for your business or scientific goal. Unlike mean regression, which targets the expected value, quantile regression targets a specific point in the conditional distribution. For example, a network provider might want to predict the 99th percentile of connection interruptions to guarantee service reliability. Once this alpha is chosen, it must remain consistent throughout the training and evaluation phases to ensure the model is optimized for the correct functional.
Practical tip
When performing hyperparameter tuning with GridSearchCV, always use the negated pinball loss (via make_scorer) to ensure the optimizer correctly identifies the best model by maximizing the score.
Select a model that minimizes pinball loss
In scikit-learn, you select an estimator that supports quantile loss. For instance, HistGradientBoostingRegressor can be configured with loss='quantile' and a specific 'quantile' parameter. Other options include QuantileRegressor. The model will attempt to minimize the pinball loss to find the value that satisfies the chosen quantile level.
from sklearn.ensemble import HistGradientBoostingRegressor
import numpy as np
X = np.random.rand(100, 1)
y = 2 * X.ravel() + np.random.normal(0, 0.5, 100)
# Target the 95th percentile
model = HistGradientBoostingRegressor(loss='quantile', quantile=0.95)
model.fit(X, y)Evaluate with mean_pinball_loss using the same alpha
To measure how well your model performs, use the mean_pinball_loss function. It is critical that the alpha parameter passed to this function is identical to the quantile targeted during model training. If you trained for the 0.95 quantile but evaluate using 0.50, the resulting error metric will be meaningless for your specific task.
from sklearn.metrics import mean_pinball_loss
y_pred = model.predict(X)
loss = mean_pinball_loss(y, y_pred, alpha=0.95)
print(f'Pinball Loss: {loss}')Create a custom scorer for tuning and validation
When using cross-validation or GridSearchCV, you cannot pass mean_pinball_loss directly because it is a loss (to be minimized) rather than a score (to be maximized). You must wrap it using make_scorer. Because scikit-learn's optimization logic expects higher values to be better, you must set greater_is_better=False. This tells the scorer to negate the loss value internally.
from sklearn.metrics import make_scorer
from sklearn.model_selection import cross_val_score
# Create a scorer for the 95th percentile
scorer = make_scorer(mean_pinball_loss, alpha=0.95, greater_is_better=False)
# Use in cross-validation
scores = cross_val_score(model, X, y, scoring=scorer, cv=5)
print(f'CV Scores: {scores}')Interpret the D² pinball score for skill assessment
The d2_pinball_score serves as a skill score, analogous to the R² coefficient for mean regression. It measures the fraction of deviance explained by your model compared to a baseline model that always predicts the alpha-quantile of the training data. A score of 1.0 indicates a perfect model, 0.0 indicates the model is no better than the baseline, and negative values indicate the model performs worse than the baseline.
from sklearn.metrics import d2_pinball_score
skill_score = d2_pinball_score(y, y_pred, alpha=0.95)
print(f'D2 Pinball Score: {skill_score}')Visualize quantile predictions correctly
Standard regression visualization involves checking if points lie on the diagonal (y_true = y_pred). However, for quantile regression, the points will not cluster on the diagonal. Instead, for a quantile alpha, you expect a specific fraction of points to fall above and below the diagonal. For example, if you predict the 0.95 quantile, approximately 95% of the actual values should be below the predicted values, assuming the model is well-calibrated.
Handle multi-output regression
If your target variable is a vector (multi-output), both mean_pinball_loss and d2_pinball_score provide a multioutput parameter. By default, this is set to 'uniform_average', which calculates the error for each output and then takes the average. You can also use 'raw_values' to get an array of errors, one for each target variable, or provide a custom array of weights to prioritize certain outputs.
Align parameter names across training and evaluation
A common source of error is the discrepancy in argument naming within scikit-learn. Estimators typically use the parameter name 'quantile' to define the target level (e.g., HistGradientBoostingRegressor(quantile=0.95)). However, the metric functions mean_pinball_loss and d2_pinball_score use the parameter name 'alpha' (e.g., mean_pinball_loss(y_true, y_pred, alpha=0.95)). Always ensure these values are synchronized to maintain mathematical consistency.
Things to check
- Verify that the estimator's quantile parameter matches the metric's alpha parameter.
- Ensure make_scorer is used with greater_is_better=False for pinball loss.
- Confirm that d2_pinball_score is used for skill assessment rather than direct error measurement.
- Check that multioutput settings are correctly configured for multi-target datasets.
Where this applies
Pinball loss is only strictly consistent for quantile prediction; it is not suitable for predicting the mean or mode. The visual interpretation of prediction error plots differs from models targeting the conditional mean.