TATECHATLAS
◎ English
Math & models

Overfitting and Learning Curves: Diagnose the Gap Without Data Leakage

Read training and validation scores together, fit preprocessing inside a pipeline and keep a final test set out of model selection.

On this page

A model that performs much better on its training data than on unseen validation data may be overfitting. The gap is a diagnostic clue, not a complete explanation: unsuitable splits, distribution changes, leakage and small samples also matter. Start with a split that represents future use, put learned preprocessing inside a pipeline and compare a suitable baseline. Learning curves then show how training and validation scores change as the training set grows. Select model complexity on development data and reserve an independent test set for the final evaluation, rather than repeatedly choosing settings against it.

Separate overfitting from other problems

For a higher-is-better score such as accuracy, strong training performance and substantially weaker validation performance suggest a generalization problem. If both scores are weak, the model or features may be inadequate instead. Do not diagnose overfitting from one number or one lucky split. Inspect sample size, labels and whether the validation examples represent deployment. A gap caused by a changed population is a different investigation from a model that memorizes details within otherwise comparable data.

Choose a split matching future use

Decide what future predictions mean: new independent examples, new customers or later observations. An ordinary random split is not suitable for every case. Repeated measurements from one subject may need grouping, and chronological prediction may need a time-aware split. Stratification preserves class proportions but does not solve dependence or time leakage. The illustrative code uses stratified folds for synthetic classification data; apply that choice to real data only when its assumptions match your intended use.

Fit preprocessing inside the pipeline

Split before fitting any transformation that learns from the data. Scaling, imputation and feature selection can leak validation information if fitted on the whole dataset. A pipeline lets cross-validation fit the scaler on each training portion and apply it to that fold's validation portion. This makes the workflow more reliable but does not repair inherently leaking features, such as a field unavailable at prediction time. Inspect feature meaning and timing as carefully as the estimator configuration.

Read the two learning curves together

A learning curve evaluates training and validation scores at several training-set sizes. Read both curves rather than just the best point. A persistent large gap can suggest high variance; weak curves close together can suggest insufficient model capacity or features. More training examples may help, but the curve does not promise a particular improvement. In the example, each row of score arrays corresponds to a training size and each column to a validation fold; averages summarize those folds.

import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, learning_curve
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

# Illustrative synthetic classification data.
X, y = make_classification(
    n_samples=500, n_features=20,
    n_informative=5, n_redundant=5, random_state=42
)
model = make_pipeline(
    StandardScaler(), LogisticRegression(max_iter=1000)
)
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
sizes, training_scores, validation_scores = learning_curve(
    model, X, y, cv=cv,
    train_sizes=np.linspace(0.2, 1.0, 5), scoring="accuracy"
)
print(sizes)
print(training_scores.mean(axis=1))
print(validation_scores.mean(axis=1))

Compare a baseline and a suitable metric

Compare against a simple baseline suited to the task before interpreting a seemingly impressive score. For imbalanced classification, high accuracy can reflect frequent-class predictions rather than useful detection of the rare class. Choose the metric to reflect the decision the model supports and keep its direction clear. The example uses accuracy only to demonstrate the learning-curve API on synthetic data. Its output is not an argument that accuracy is the right metric for every real classification problem.

Change complexity and regularization deliberately

Try a small, planned set of complexity or regularization changes and evaluate them using the same development protocol. With LogisticRegression, smaller C means stronger regularization; max_iter controls the optimization iteration limit, not regularization strength. Scaling and convergence should also be checked. Change one meaningful factor at a time and compare validation behavior rather than trying to maximize the training score. A simpler model can generalize better, but that must be judged on representative validation data.

Inspect variation across validation folds

Inspect individual fold scores as well as their mean. A wide spread can indicate that the estimate is sensitive to the selected examples or that folds differ in difficulty. Their standard deviation describes variation between these folds; it is not automatically a confidence interval. Cross-validation splits share training observations, and their scores are not independent experimental replicates. Use the spread as a reason to inspect the protocol, sample and groups, rather than as proof of a precise population error bound.

Keep a final evaluation outside selection

Freeze the chosen features and parameters before the final held-out evaluation. Repeatedly inspecting that test set while choosing settings makes it part of selection and weakens its independence. Record the dataset version, feature definitions, split rules, random state, metric and software versions. The code creates synthetic data and prints computed score summaries; no particular numeric score or benchmark is claimed here. A reproducible workflow makes later changes understandable and helps separate a genuine modeling improvement from a changed evaluation procedure.

Things to check

  • Use validation splits that represent future cases, groups or time periods.
  • Fit learned preprocessing on training portions only, preferably inside a pipeline.
  • Read training and validation curves together with baseline performance and fold variation.
  • Choose parameters on development data and keep final test observations out of selection.

Learning curves depend on sample quality, split design, model and metric. They suggest possible causes but do not prove that more data or stronger regularization will solve the task. The synthetic example is illustrative and was not executed as a benchmark.

Sources

  1. scikit-learn: learning and validation curves ↗
  2. scikit-learn: learning_curve API ↗
  3. scikit-learn: data leakage and pipelines ↗
  4. scikit-learn: cross-validation ↗
  5. scikit-learn: LogisticRegression ↗
Back to top ↑