Validating Forecasts with TimeSeriesSplit and Per-Horizon Error Analysis
Use TimeSeriesSplit to evaluate time-ordered forecasts without leakage, compute per-horizon errors, and choose a strictly consistent metric that matches the decision cost.
On this page
Key idea
Use TimeSeriesSplit to generate time-ordered train/test folds so every evaluation uses only past data, then compute a regression metric separately for each forecast horizon and compare those per-horizon errors. The final metric is not a single global average; it is the decision-relevant summary of how error changes as the horizon grows. Choose a strictly consistent scoring function, such as mean_absolute_error for a median target or mean_squared_error for a mean target, and use it both for training loss and evaluation. This alignment ensures the metric truthfully reflects the cost of the chosen point forecast. For example, with 12 monthly samples, TimeSeriesSplit(n_splits=3, test_size=2) creates three folds where each test set covers two months and the training set only contains earlier months. If the per-horizon MAE rises from 1.2 at horizon 1 to 2.8 at horizon 3, the business cost of a three-month-ahead decision is materially higher than a one-month-ahead decision. You can then decide whether to accept the degradation, add features, or restrict the forecast horizon. The approach assumes equally spaced samples and a point forecast, not a full probabilistic distribution.
Validating Forecasts with TimeSeriesSplit
TimeSeriesSplit is a cross-validator designed for time-ordered data. It prevents data leakage by ensuring that the model is trained only on past observations and evaluated on subsequent time-ordered folds. Standard cross-validation methods are inappropriate for time series because they can train on future data and evaluate on past data. TimeSeriesSplit returns train and test indices so that each test set covers a later time window than its training set. The training set accumulates data from previous splits, and successive training sets are supersets of earlier ones. This structure matches the temporal order of forecasting, where future values must never be used to predict the past.
To ensure comparable metrics across folds, samples must be equally spaced. Once this condition is met, each test set covers the same time duration while the training set grows. The class accepts parameters such as n_splits, max_train_size, test_size, and gap. The default test size is n_samples // (n_splits + 1), and the gap parameter excludes samples from the end of each training set before the test set. The split method yields train and test index arrays that can be used to slice the data. This makes it straightforward to loop over folds, fit a model, and evaluate it on each test window.
For a fixed test_size, the folds produce test sets of equal length. Each fold evaluates the model on a different horizon, and the per-fold error can be recorded separately. This is the foundation for analyzing error across forecast horizons. The same evaluation loop can be extended to store predictions for each horizon so that metrics are computed per horizon rather than only per fold. The key requirement is that the data remain time-ordered and that no future information leaks into the training step.
import numpy as np
from sklearn.model_selection import TimeSeriesSplit
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error
X = np.arange(12).reshape(-1, 1)
y = np.array([1.0, 1.8, 3.2, 4.1, 5.0, 5.9, 6.8, 7.7, 8.6, 9.5, 10.4, 11.3])
tscv = TimeSeriesSplit(n_splits=3, test_size=2)
for fold, (train_idx, test_idx) in enumerate(tscv.split(X), start=1):
X_train, X_test = X[train_idx], X[test_idx]
y_train, y_test = y[train_idx], y[test_idx]
model = LinearRegression()
model.fit(X_train, y_train)
pred = model.predict(X_test)
print(f"Fold {fold}: test indices {test_idx}, MAE {mean_absolute_error(y_test, pred):.3f}")
Analyzing Error Across Forecast Horizons
After obtaining predictions from each fold, the next step is to compare metrics on each forecast horizon. A single global metric hides how prediction quality degrades as the time gap from the training set increases. By grouping errors by horizon, you can see whether short-horizon forecasts are accurate and long-horizon forecasts deteriorate, or whether the error pattern is irregular. This comparison is essential for deciding how far ahead a forecast can be trusted for a specific business action.
The per-horizon error is computed by selecting all predictions whose horizon matches a given value and applying the chosen regression metric to the corresponding true values. This requires storing the horizon index for each prediction, which is natural when forecasting multiple steps ahead from a single model or when evaluating each fold at a specific horizon. The resulting table or plot shows the error trajectory across horizons. It directly answers the question of how much worse the forecast becomes as the prediction window lengthens.
This analysis also helps separate structural error from noise. If error increases steadily with horizon, the model may lack the information needed for longer-term predictions. If error spikes at a particular horizon, that may indicate a structural break, seasonality, or insufficient training data for that range. The per-horizon view is more informative than a single aggregate score because it reveals where the forecast loses value. It also supports the next step of linking the metric to decision costs, because each horizon can be mapped to a specific action or risk.
from sklearn.metrics import mean_absolute_error
# y_true and y_pred are aligned per horizon
horizon_errors = {}
for horizon in range(1, max_horizon + 1):
mask = (horizon_index == horizon)
horizon_errors[horizon] = mean_absolute_error(y_true[mask], y_pred[mask])
Linking Metrics to Decision Costs
The choice of scoring function should start from the ultimate goal and application of the prediction, not from convenience. It is useful to distinguish two steps: predicting and decision making. Predicting involves choosing a property or functional of the response distribution, such as the mean, median, or a quantile. Decision making involves acting on that prediction. The metric used for evaluation should be strictly consistent with the target functional chosen during prediction. A strictly consistent scoring function guarantees that truth telling is an optimal strategy in expectation, which means the metric accurately reflects the distance between predictions and true targets.
For regressors, the typical target functionals are the mean and the median. Mean squared error is aligned with the mean, while mean absolute error is aligned with the median. Using a metric that is not consistent with the target functional can produce misleading comparisons. For example, if the model is trained to predict the median but evaluated with squared error, the evaluation may not reflect the actual cost of the chosen point forecast. The scikit-learn documentation recommends using a strictly consistent scoring function for the chosen functional and using it both as the loss function for training and as the metric for evaluation and model comparison.
Once the metric is chosen, it should be connected to the business cost of the decision. The final metric is not just a number; it is a proxy for the cost of forecast errors. If the per-horizon MAE is 1.2 at horizon 1 and 2.8 at horizon 3, the cost of a three-month-ahead decision is more than double that of a one-month-ahead decision. This comparison can justify restricting the forecast horizon, adding features, or accepting the degradation. The metric must be interpreted in the context of the application, because the same error value can have different consequences in different domains.
from sklearn.metrics import mean_absolute_error, mean_squared_error
# Strictly consistent scoring functions for point forecasts
mae = mean_absolute_error(y_true, y_pred) # aligns with median target
mse = mean_squared_error(y_true, y_pred) # aligns with mean target
Selecting Consistent Scoring Functions
A strictly consistent scoring function should be selected so that it aligns with the target functional of the point forecast. For a mean target, mean squared error is the standard choice. For a median target, mean absolute error is the standard choice. These functions guarantee that the expected score is minimized when the prediction equals the target functional, which makes the metric a truthful measure of forecast quality. The scikit-learn documentation emphasizes that once a strictly consistent scoring function is chosen, it is best used both as the loss function for model training and as the metric for model evaluation and comparison.
The table of available scorers in the scikit-learn model evaluation documentation includes regression metrics such as mean_absolute_error, mean_squared_error, root_mean_squared_error, and median_absolute_error. Each has a specific alignment with a target functional. The choice should be driven by the business objective and the statistical property of the forecast that matters. If the goal is to minimize average absolute deviation, mean absolute error is appropriate. If the goal is to minimize average squared deviation, mean squared error is appropriate. The metric must be consistent with the prediction target to avoid biased evaluation.
Using the same metric for training and evaluation ensures that the model optimizes the quantity that will be measured. This alignment prevents situations where a model performs well on an inconsistent metric but poorly on the actual decision-relevant quantity. The per-horizon analysis then becomes meaningful because the metric reflects the true cost of the forecast at each horizon. The final decision can be based on the error trajectory and the business impact of errors at different horizons, rather than on a single aggregate score that may obscure important patterns.
from sklearn.metrics import mean_absolute_error
# Example: 12 monthly samples, 3 folds, test_size=2
# Fold 1 test indices: 6,7; Fold 2 test indices: 8,9; Fold 3 test indices: 10,11
# Per-horizon MAE: horizon 1 = 1.2, horizon 2 = 1.9, horizon 3 = 2.8
# Decision: restrict forecast to horizon 2 or improve long-horizon features
Applicability
- Are the samples equally spaced so that each test set covers the same time duration?
- Is every test set strictly after its training set, with no future data used for training?
- Is the chosen metric strictly consistent with the target functional, such as mean or median?
- Are per-horizon errors computed separately instead of collapsing everything into one global number?
- Does the final metric connect to the actual business cost of the decision?
- Is the forecast a point forecast rather than a full probabilistic distribution?
- Are the training and test sets disjoint in time for every fold?
- Is the metric used consistently for both model training and evaluation?
Where this applies
TimeSeriesSplit requires samples to be equally spaced to ensure comparable metrics across folds. The approach assumes a point forecast rather than a full probabilistic distribution. It does not automatically account for asymmetric business costs or nonstationarity in the error structure. Per-horizon error comparisons can be noisy when the number of folds is small, and the final summary metric must be interpreted against the specific decision context.