Compare a time-series forecast with a seasonal baseline
Build a previous-season reference, evaluate it on later observations and distinguish calendar alignment from leakage and missing data.
On this page
The short answer
A seasonal baseline predicts an observation using the value from the corresponding earlier season. For regularly spaced monthly observations with annual seasonality, the reference is twelve observations earlier. Compare your model and baseline on the same later dates using only information available at each forecast origin. A low training error alone does not establish that a model improves on this simple reference.
Choose a reference that matches the task
An overall mean ignores where a value lies within a repeating cycle. When demand reliably differs between months, last year’s corresponding month can be a more informative reference. This is a hypothesis to evaluate, not proof that every series is seasonal.
Choose the period from the observation schedule and the actual process. Twelve rows mean a year only for a complete monthly series. Twelve rows in daily data or a table with missing months describe a different interval. A baseline should match the forecast target rather than a convenient array length.
Keep the time index regular and explicit
Sort observations chronologically and inspect duplicate timestamps, gaps and units. In pandas, shifting values by periods is different from shifting the index with a frequency. For a lagged reference, values must be aligned with the future dates they predict.
Do not silently fill gaps using later values, because that can reveal information unavailable at the forecast origin. If the series is irregular, explicitly align matching calendar periods or choose another reference. Decisions about missing observations should be applied consistently to the model and baseline.
A small monthly example
The example uses Python’s standard library and twenty-four synthetic, equally spaced monthly observations. The first twelve values represent one year, and the next twelve are two units higher. They are illustrative data, not measurements or evidence of real forecasting performance.
For the second year, a twelve-observation lag predicts the first year’s values. The first twelve positions have no earlier season in this dataset, so they cannot be scored by this baseline. Never replace those absent references with future observations.
# Illustrative monthly observations, January through December twice.
values = [10, 12, 15, 18, 20, 22, 21, 19, 17, 14, 12, 11]
values += [value + 2 for value in values]
period = 12
actual = values[period:]
predicted = values[:-period]
mae = sum(abs(a - p) for a, p in zip(actual, predicted)) / len(actual)
print(predicted[:3])
print(mae)
# Expected illustrative output:
# [10, 12, 15]
# 2.0Interpret the error in the original units
Here every forecast misses by two units, so the illustrative mean absolute error is 2.0. This arithmetic describes the supplied toy data only. It says nothing about an unseen real series or whether two units is acceptable for the application.
Evaluate your model on those same target dates and in the same units. Compare errors by season as well as overall: an average can hide repeated failures during a critical month. Document which observations were excluded because a valid target or lagged reference was unavailable.
Split past from future
A random train/test split can put future observations into training while earlier dates appear in evaluation. For forecasting, use chronological evaluation. At each origin, construct features and fit any preprocessing using only information available by that origin.
TimeSeriesSplit provides successive time-ordered training and test windows with growing training sets by default. Its comparable-duration interpretation assumes equally spaced observations. Choose test_size and any gap to match the task; the library does not determine the business forecasting horizon for you.
Distinguish one-step and multi-step forecasts
In rolling evaluation, a value that was unavailable at one origin may become known before the next origin. A forecast issued for several future months at once has a different information set. Do not evaluate that multi-step forecast as though each intervening actual value had already arrived.
A seasonal lag is usable only when the referenced observation is known at the relevant origin. For horizons longer than one seasonal period, some direct lag references point into the forecast interval and need a defined strategy. State that strategy before comparing errors rather than selecting it after seeing the targets.
Check why the baseline succeeds or fails
A change in level, trend, a moved holiday or a changed business process can make last season a poor predictor. Inspect dated residuals and known events. Beating the baseline during one unusual window is weaker evidence than improving across several representative later windows.
Compare with a last-observation or other simple baseline when the seasonal reference performs badly. Keep the evaluation dates identical. Avoid tuning many models against a final held-out period until it effectively becomes another training set.
Decide whether the added model is useful
A more complex model is useful when it improves a relevant error measure on later data enough to justify its maintenance and operating cost. Specify what matters: typical absolute error, large misses, particular seasons or the cost of overprediction versus underprediction.
Retain the baseline as a reference after deployment. Changes in the data process can make an earlier advantage disappear. Monitoring comparable errors supports a concrete decision to keep, revise or simplify the model without inventing a performance guarantee.
Things to check
- Confirm a regular time index and a justified seasonal period.
- Use only observations known at each forecast origin.
- Score model and baseline on identical later targets.
- State the forecast horizon and missing-data policy.
Where this applies
A seasonal baseline needs earlier matching observations and sufficiently stable repetition. It does not establish causal relationships, guarantee accuracy or automatically handle structural breaks, irregular timestamps or multi-step horizons.