TATECHATLAS
◎ English
Math & models

Validate a time-series forecast without leaking the future

Use chronological splits and fit preprocessing only on past data.

On this page

For a forecasting test, train on earlier observations and evaluate on later ones. Random splitting can expose the model to information it would not have when making a real forecast.

Start with the moment when a forecast must be made

Suppose a shop predicts the next seven days of demand each Sunday evening. The evaluation must imitate that moment: next week’s sales and information published after Sunday cannot be inputs. A random split mixes earlier and later records and may make a model look more useful than it will be in operation.

Write down the forecast origin, horizon and when each input becomes available. Calendar features can be known ahead; an observed temperature or a revised sales total may not be. For weather, use the forecast available at that time, not the later observation, if that matches the intended application.

Keep the forecast horizon explicit

Set the test window to match the horizon you need. In each split, fit scaling and other learned preprocessing on training data only. A rolling feature must use information available at prediction time.

from sklearn.model_selection import TimeSeriesSplit

observations = list(range(12))
splitter = TimeSeriesSplit(n_splits=3, test_size=2, gap=1)
for train, test in splitter.split(observations):
    print(train.tolist(), test.tolist())

Read a chronological split in concrete terms

The small example above uses 12 ordered observations, three evaluation windows of two observations and a one-row gap. Its first fold trains on positions 0–4 and evaluates on 6–7; the next trains on 0–6 and evaluates on 8–9; the last trains on 0–8 and evaluates on 10–11.

The gap represents a chosen exclusion interval, not a general cure for leakage. Select it according to publication delays or overlapping target windows. TimeSeriesSplit counts rows, so two rows do not necessarily mean two days. Sort timestamps and check duplicates, missing periods and spacing before interpreting the windows.

Check temporal assumptions

Check sorting, missing dates and spacing. Use a gap where the information delay or overlapping labels require one. Keep a final untouched period for assessment.

Look for leakage outside the split itself

Fit scalers, imputers and feature selection only on each training fold; apply those fitted transformations to its later test fold. A Pipeline helps keep learned preprocessing together, but cannot repair a feature that already contains future information.

For daily sales, a trailing mean for a prediction at day t should use earlier observations, for example shift(1).rolling(7).mean(). A centered window can include future days. When predicting a whole week at once, even actual sales from its first days are unavailable at the forecast origin; do not feed them into later days as if they were known.

Compare against a simple forecast before tuning

Use a baseline matching the task: the latest known value or the corresponding previous seasonal period. Compute errors for the same forecast origins and horizon as the candidate model. In the illustrative calculation below, actual values 100, 110 and 90 with predictions of 100 give MAE ≈ 6.67 in the target’s units.

Inspect errors by horizon and time period; a single average may hide poor performance on busy days. Percentage metrics can behave badly near zero. Reserve a final period that does not guide model selection, and recheck performance when the data-generating conditions change.

actual = [100, 110, 90]
predicted = [100, 100, 100]
mae = sum(abs(a - p) for a, p in zip(actual, predicted)) / len(actual)
print(round(mae, 2))

Things to check

  • Split chronologically.
  • Fit preprocessing within each training fold.
  • Check feature availability at forecast time.

This example assumes chronologically ordered observations. TimeSeriesSplit uses row positions; verify equal spacing for comparable time windows.

Sources

  1. scikit-learn: TimeSeriesSplit ↗
  2. scikit-learn: common pitfalls ↗
  3. pandas: time series ↗
  4. pandas: rolling windows ↗
Back to top ↑