TATECHATLAS
◎ English
Math & models

Evaluating prediction intervals with empirical coverage and width

Prediction intervals aim to contain future individual observations, while confidence intervals target a mean or other parameter, and the two are not interchangeable. Empirical coverage is the held-out fraction inside the interval and width is the distance between quantile limits; a constructed five-value example gives 80 percent coverage and mean width 1.8.

On this page

A prediction interval is a range intended to contain a future individual observation with a stated probability, while a confidence interval is a range for an unknown parameter such as a conditional mean. The two are not interchangeable because they answer different questions and have different widths. Empirical coverage is the fraction of held-out observations that fall inside the predicted interval, and interval width is the distance between the lower and upper quantile limits. A small illustrative sample with realized values [2,5,8,4,10] and intervals [(1,3),(4,6),(6,7),(3,5),(9,11)] gives four inside and one outside, so empirical coverage is 4/5 = 80 percent, and the widths [2,2,1,2,2] average 1.8. This constructed sample is only an illustration of the arithmetic, not a calibration study and not a scikit-learn benchmark. In scikit-learn, quantile regression can produce prediction intervals by fitting separate models at complementary quantile levels, for example alpha=0.05 and alpha=0.95 to form a 90 percent interval. The scikit-learn quantile regression example shows that the 5 percent and 95 percent quantile models are used to define a central interval, and the gradient boosting quantile example shows the same idea with GradientBoostingRegressor using loss='quantile' and alpha set to 0.05, 0.5, and 0.95. The source also shows that calibration must be checked on held-out data because training coverage can be close to the nominal level while test coverage can be lower. In the gradient boosting example, training coverage is about 0.903 while test coverage is about 0.796, and a bootstrap confidence interval for test coverage is reported as lying between 73.8 percent and 84.4 percent. That gap illustrates that nominal coverage is not an individual guarantee and that interval width can be too narrow on new data. Quantile models can also cross, meaning the lower quantile prediction can exceed the upper one for some inputs, which breaks the interval interpretation. Coverage can also vary across subgroups, so a single overall coverage number can hide poor performance in important segments. When choosing the nominal level, decide whether the goal is a central interval such as 90 percent or a one-sided bound, and remember that the nominal level refers to long-run frequency over many future observations, not to a promise about any single case. To count held-out coverage, compare each held-out observation with its predicted lower and upper limits and compute the proportion where the observation lies between them. To measure width, subtract the lower limit from the upper limit for each case and summarize with the mean or a percentile, while noting that narrower intervals are only better if coverage remains acceptable. A useful practical comparison is to look at width at a coverage level that is actually achieved on held-out data, because an overly narrow interval that misses its nominal target is not useful even if it looks efficient. Validation must keep training and evaluation separate to avoid leakage, and quantile estimation for extreme levels can be unstable because very few observations inform the tails. The main limitation is that empirical coverage on a small held-out set is noisy, that quantile crossing and subgroup variation can distort interpretation, and that the illustrative numbers here are hypothetical and not observed results from an executed pipeline.

State what the interval predicts

A prediction interval is meant to contain a future individual observation with a stated probability, while a confidence interval is meant to contain an unknown parameter such as a conditional mean. That distinction matters because the two intervals answer different questions and are usually different widths. The prediction interval must account for both uncertainty about the underlying signal and the variability of a new observation around that signal, whereas a confidence interval for a mean only reflects uncertainty about the estimated center. In the scikit-learn quantile regression example, the 5 percent and 95 percent quantile models are used to form a central interval for observations, and the documentation frames this as creating prediction intervals rather than intervals for a mean. The gradient boosting quantile example makes the same point by training separate quantile regressors and using the 0.05 and 0.95 models to define a predicted 90 percent interval. The illustrative values [2,5,8,4,10] with intervals [(1,3),(4,6),(6,7),(3,5),(9,11)] are used here only to show the arithmetic of checking whether an observation falls inside a predicted range. They are not evidence about any fitted model.

Choose the nominal level

The nominal level is the target long-run coverage probability, such as 90 percent for a central interval formed by the 0.05 and 0.95 quantiles. Choosing it depends on the decision problem, not on the data alone. A central interval is common when you want to bound most future observations, while a one-sided bound may be more appropriate when only an upper or lower limit matters. The scikit-learn gradient boosting quantile example explicitly says that the models for alpha=0.05 and alpha=0.95 produce a 90 percent coverage interval because 95 percent minus 5 percent equals 90 percent. The quantile regression example similarly uses 0.05 and 0.95 to define a central 90 percent interval and to identify observations outside that central region. The key caution is that nominal coverage is not an individual guarantee. A 90 percent interval does not mean that any specific observation has a 90 percent chance of being inside the computed range in a way that overrides the model's assumptions and calibration. The illustrative sample uses a nominal target that is not stated as a fitted model property; the 80 percent empirical coverage there is only the result of counting four inside and one outside in a tiny constructed set.

Count held-out coverage

Held-out coverage is the proportion of previously unseen observations that fall between the predicted lower and upper limits. To compute it, compare each held-out observation with its corresponding interval and count the cases where the observation is greater than or equal to the lower limit and less than or equal to the upper limit, then divide by the number of held-out cases. The gradient boosting quantile example implements this idea directly with a coverage_fraction function that returns the mean of the logical AND of the two comparison conditions. In that example, training coverage is reported as about 0.903, while test coverage is reported as about 0.796, and a bootstrap confidence interval for test coverage is reported as lying between 73.8 percent and 84.4 percent. Those numbers show that calibration can look good on training data and still be too optimistic on held-out data. The illustrative sample follows the same counting logic: four of the five values fall inside their intervals, so empirical coverage is 4/5, or 80 percent. That calculation is only a demonstration of the definition, not a model evaluation result.

Measure interval width

Interval width is usually measured as the difference between the upper and lower quantile limits for each case. A summary such as the mean width or a high percentile of widths can be used to compare methods, but width should be judged together with coverage rather than alone. Narrower intervals are only better if they still achieve acceptable coverage on held-out data. In the illustrative sample, the widths are [2,2,1,2,2], and the mean width is 1.8. That number is computed directly from the five interval pairs and is independent of any model fit. The scikit-learn examples do not report a single width number in the same way, but they do show that the distance between the 0.05 and 0.95 predictions defines the predicted interval and that this distance can change with the input, especially under heteroscedastic noise. The quantile regression example notes that with increasing noise variance the interval between the 5 percent and 95 percent quantiles becomes wider as x increases. That behavior is one reason why a single average width can be misleading if the interval width varies a lot across the input space.

Work through five observations

The illustrative sample contains five realized values [2,5,8,4,10] and five corresponding intervals [(1,3),(4,6),(6,7),(3,5),(9,11)]. Checking each case: the first value 2 is inside (1,3); the second value 5 is inside (4,6); the third value 8 is outside (6,7); the fourth value 4 is inside (3,5); and the fifth value 10 is inside (9,11). That yields four inside and one outside, so empirical coverage is 4/5 = 0.80, or 80 percent. The widths are 3 minus 1 equals 2, 6 minus 4 equals 2, 7 minus 6 equals 1, 5 minus 3 equals 2, and 11 minus 9 equals 2, so the width list is [2,2,1,2,2] and the mean width is 1.8. This is a fully constructed example, so it should be read as a definition check rather than as a claim about a real model or dataset. It also shows why a single coverage number can hide detail: one miss in five observations changes the result by 20 percentage points, which is a large swing for such a small sample.

Compare width at useful coverage

A practical comparison asks whether an interval is narrow enough to be useful while still achieving coverage that is actually observed on held-out data. An interval that is too narrow can look efficient but fail its coverage target, while an interval that is very wide may achieve coverage easily but provide little guidance. The gradient boosting quantile example illustrates this trade-off by showing that the tuned models improve the interval shape but still produce test coverage below the nominal 90 percent, with a bootstrap interval for test coverage between 73.8 percent and 84.4 percent. That means the interval is not reliably achieving its nominal goal on held-out data, even though training coverage is close to 90 percent. When comparing methods, it is therefore more informative to report both coverage and width on the same held-out set, and to check whether the achieved coverage is acceptable for the intended use. The illustrative sample is too small to support a meaningful width-versus-coverage comparison, but it does show the basic idea: the mean width is 1.8 and the empirical coverage is 80 percent, so any claim that the intervals are good should be judged against a nominal target and a larger held-out set.

Check validation and leakage

Coverage and width must be evaluated on data that were not used to fit the quantile models, because using the training set can make calibration look better than it really is. Leakage can occur if the same observations influence both the fitted quantile curves and the evaluation, or if the evaluation uses information that would not be available at prediction time. The gradient boosting quantile example separates training and test data and reports different coverage values for each, which is the right structure for checking calibration. It also uses cross-validated tuning for the 0.05 quantile model and reports that the tuned models improve the 90 percent interval, but it still finds that test coverage is below the nominal level. The quantile regression example similarly evaluates out-of-sample performance with cross-validation when comparing QuantileRegressor and LinearRegression. The cautionary point is that quantile estimates for extreme levels can be unstable because they depend on relatively few observations, and that instability can be amplified if the evaluation is not properly separated from training. The illustrative sample does not involve fitting or leakage, but it still shows why a small held-out set can produce volatile coverage estimates.

Explain calibration limits

Calibration here means that the fraction of held-out observations inside the predicted interval is close to the nominal level on average. Good calibration on one dataset does not guarantee good calibration on another, and average calibration can hide variation across subgroups or across regions of the input space. The gradient boosting quantile example reports training coverage near 0.903 and test coverage near 0.796, and it explicitly says the estimated interval is too narrow to cover 90 percent of the test points. It also reports a bootstrap confidence interval for test coverage between 73.8 percent and 84.4 percent, which quantifies the uncertainty in the held-out estimate. Two further limits are worth stating. First, quantile predictions can cross, so the lower prediction can exceed the upper prediction for some inputs, which breaks the interval interpretation. Second, coverage can vary across subgroups, so a single overall coverage number may not reflect performance where it matters most. The illustrative sample does not exhibit crossing or subgroup analysis, but it does show that empirical coverage is just a proportion from a finite set and should not be treated as a precise long-run guarantee.

Things to check

  • Confirm whether the interval is for a future observation or for a mean parameter before interpreting coverage.
  • Verify that lower and upper quantile predictions are computed from separate models or a method that preserves order.
  • Check that coverage is computed on held-out data, not on the training set used to fit the quantile models.
  • Inspect whether quantile predictions cross for any inputs, which would invalidate the interval semantics.
  • Review whether coverage is stable across meaningful subgroups rather than relying only on a single overall fraction.
  • Treat the five-value example as arithmetic illustration only, not as evidence of model quality or scikit-learn performance.

The five-value example is constructed for illustration and does not constitute a calibration study or a benchmark result. Empirical coverage and width computed from a small held-out set are noisy and can mislead if presented as precise estimates. Nominal coverage describes long-run frequency over many future observations, not a guarantee for any single observation. Quantile models can cross, and coverage can vary across subgroups, so a single overall number may hide poor performance in specific segments. The scikit-learn excerpts describe methods and show example coverage numbers, but the concrete test-set coverage values there come from that example's synthetic setup and modeling choices, not from a universal property of quantile regression. Extreme quantiles are estimated from relatively few observations and can be unstable. This response uses only the supplied primary sources and the specified illustrative numbers; it does not add executed tests or observed retrieval results.

Sources

  1. scikit-learn: prediction intervals ↗
  2. scikit-learn: quantile regression ↗
Back to top ↑