TATECHATLAS
◎ English
Artificial intelligence / Tip

Fix the evaluation set and comparison rules before comparing models

Fix the evaluation set, metrics, aggregation rules, and comparison rules before running experiments so results stay comparable and per-subset failures are not hidden by one aggregate number.

On this page

A single final metric collapses the experiment into one number and hides where the model wins or loses. Fix the evaluation set, the metrics, the aggregation order, and the comparison rules before running the experiment. The set must state the splits, subsets, weights, metric directions, and the aggregation method. The rules must state that models are compared on the same splits, that each subset is reported separately, and that the aggregate is a weighted average. This makes results reproducible across runs and authors. The Hugging Face Evaluate library and scikit-learn scoring API both support this workflow. Evaluate lets you load metrics and compute them consistently, while scikit-learn's scoring parameter defines evaluation rules for model selection. The key is to record the spec in a file or configuration so that the same inputs always produce the same outputs. Without this, a model can appear better because the aggregate hides a failure on a critical subset. The example shows two classifiers with accuracy 0.91 versus 0.89, but on subset B recall drops from 0.78 to 0.62. If only the aggregate is reported, the loss is invisible. With the spec fixed, the report shows both the aggregate and the per-subset breakdown, so the comparison is transparent and defensible.

Context and recommendation

The evaluation set must be fixed before any model is compared. This means recording the splits, subsets, weights, metrics, and their directions in a file or configuration. The Hugging Face Evaluate library provides tools to load metrics and compute them consistently, while scikit-learn's scoring parameter defines evaluation rules for model selection. Both support reproducible evaluation when the rules are written down. The recommendation is to create an evaluation specification that states exactly which examples are used, which metrics are computed, and how the results are aggregated. Without this, two runs or two authors can produce different numbers that are not comparable. The spec should include the order of aggregation, such as weighted average over subsets, and the rule that models are compared on the same splits. This prevents accidental changes in preprocessing or data ordering from changing the conclusion. The goal is to make the experiment reproducible and the comparison transparent.

The evaluation spec should be stored in a file such as YAML or JSON so that it can be versioned and reused. The file should list the splits, subsets, weights, metrics, and comparison rules. For example, a classification experiment might use two subsets, A and B, with weights 0.6 and 0.4, and two metrics, accuracy and recall, both maximizing. The comparison rule states that the same splits are used, that each subset is reported separately, and that the aggregate is a weighted average. This file becomes the single source of truth for the experiment. When two models are run, they both read the same spec, so the results are directly comparable. The report should include the per-subset scores and the aggregate, not only the aggregate. This way, a model that wins on average can be seen to lose on a critical subset. The spec also records the metric directions, such as accuracy and recall maximizing, so that the aggregation is meaningful.

The Hugging Face Evaluate library and scikit-learn scoring API are practical tools for this workflow. Evaluate allows loading metrics and computing them on datasets in a consistent way. scikit-learn's scoring parameter defines model evaluation rules for cross-validation and parameter search. Both libraries support the idea that evaluation is a set of rules, not just a single number. The key is to record the rules in a file or configuration so that the same inputs always produce the same outputs. This is especially important when multiple authors or multiple runs are involved. The spec should be shared with the team and used as the basis for the final report. The report should show the raw per-subset numbers, the aggregate, and the comparison rules. This makes the experiment auditable and the conclusions defensible.

The example in the task shows two classifiers with accuracy 0.91 versus 0.89, but on subset B recall drops from 0.78 to 0.62. If only the aggregate is reported, the loss is invisible. With the spec fixed, the report shows both the aggregate and the per-subset breakdown, so the comparison is transparent and defensible. The same principle applies to any domain where multiple subsets or metrics are used. The evaluation set must be fixed, the rules must be recorded, and the report must show the breakdown. This prevents the common mistake of choosing a model based on one number that hides important failures. The spec should be created before the experiment starts, not after the results are known. This ensures that the comparison is fair and that the conclusions are based on the pre-registered rules.

eval_spec:
  splits: [train, validation, test]
  subsets:
    A: {weight: 0.6}
    B: {weight: 0.4}
  metrics:
    accuracy:
      direction: maximize
    recall:
      direction: maximize
  comparison:
    same_splits: true
    report_per_subset: true
    aggregate: weighted_average

Rationale and trade-offs

A single final metric simplifies ranking but hides the distribution of errors across subgroups, horizons, and error types.

The trade-off is clear: recording rules requires extra effort up front, but a single aggregate can reverse the ranking when a critical subset is considered. Writing a spec file takes minutes; discovering after deployment that a model fails on a minority subgroup takes weeks. The cost asymmetry favors the spec.

In the example, Model X has accuracy 0.91 and Model Y has accuracy 0.89. On subset B, recall drops from 0.78 for Model Y to 0.62 for Model X. If the aggregate is a weighted average of accuracy across subsets A and B with weights 0.6 and 0.4, Model X scores 0.91 overall while Model Y scores 0.89. The aggregate difference is only 0.02, yet the recall gap on subset B is 0.16, a reversal that the single number conceals.

The fixed spec forces the report to surface both the aggregate and the per-subset breakdown. Without the spec, an analyst might pick Model X based on the higher aggregate and miss the recall failure on subset B. With the spec, the comparison is transparent and the failure is visible before deployment.

# Evaluation spec used for model comparison
spec = {
    "splits": ["train", "validation", "test"],
    "subsets": {"A": 0.6, "B": 0.4},
    "metrics": {"accuracy": "maximize", "recall": "maximize"},
    "aggregate": "weighted_average"
}

# Example: compare two models on the same splits with per-subset reporting
# Model X: accuracy 0.91, recall on subset B 0.62
# Model Y: accuracy 0.89, recall on subset B 0.78
# The aggregate alone hides the subset B failure for Model X

# The same spec is used for both models so the comparison is fair
# Report per-subset scores and the weighted aggregate
# This prevents a high aggregate from hiding a low score on a critical subset

Concrete illustration

Consider two classifiers evaluated on a dataset split into subset A (60 percent of examples) and subset B (40 percent of examples). Model X achieves accuracy 0.91 on A and 0.90 on B. Model Y achieves accuracy 0.88 on A and 0.90 on B. The weighted aggregate accuracy is 0.6 times 0.91 plus 0.4 times 0.90 equals 0.906 for Model X, and 0.6 times 0.88 plus 0.4 times 0.90 equals 0.888 for Model Y. Model X wins by 0.018 on the aggregate.

Now examine recall on subset B specifically. Model X has recall 0.62 on B while Model Y has recall 0.78 on B. The gap is 0.16, far larger than the aggregate accuracy difference of 0.018. If subset B represents a critical minority group where missed positives carry high cost, Model X's apparent advantage is illusory.

The fixed spec states that recall on subset B must be reported separately and that the comparison rule requires per-subset visibility. When the report is generated from the spec, both the aggregate and the per-subset recall appear side by side. The reader sees that Model X wins on aggregate accuracy but loses badly on subset B recall, and can make an informed decision rather than trusting a single number.

This worked example shows why the spec matters operationally. Without it, the analyst computes one accuracy number, sees 0.91 versus 0.89, and selects Model X. With it, the same analyst is required to show the breakdown, and the recall failure on subset B becomes impossible to overlook.

eval_spec:
  splits: [train, validation, test]
  subsets:
    A: {weight: 0.6}
    B: {weight: 0.4}
  metrics:
    accuracy:
      direction: maximize
    recall:
      direction: maximize
  comparison:
    same_splits: true
    report_per_subset: true
    aggregate: weighted_average

# Model X: accuracy 0.91, recall on subset B 0.62
# Model Y: accuracy 0.89, recall on subset B 0.78
# Aggregate alone hides the subset B failure for Model X
# The report must show per-subset scores and the weighted aggregate

Limits of applicability

This approach does not replace a separate final test set when hyperparameters are tuned. It does not by itself provide statistical significance; a formal test is needed for small or noisy differences. Results remain comparable only if the evaluation set and rules stay unchanged, because changing splits or weights requires a new comparison. It also does not remove the need to choose the validation strategy, such as cross-validation or chronological splits, based on the data type. The method assumes multiple subsets or metrics exist and that the decision is based on model comparison. It cannot guarantee that a weighted aggregate is the right business objective, and it does not align embedding spaces or establish relevance ground truth without paired encoders and independent reference labels.

The approach is most useful when the evaluation set contains distinct subgroups with different operational costs, when multiple authors or teams must reproduce results, and when the decision threshold is close enough that a single number could flip the ranking.

eval_spec:
  splits: [train, validation, test]
  subsets:
    A: {weight: 0.6}
    B: {weight: 0.4}
  metrics:
    accuracy:
      direction: maximize
    recall:
      direction: maximize
  comparison:
    same_splits: true
    report_per_subset: true
    aggregate: weighted_average

Applicability

  • Does the evaluation spec list every split, subset, weight, and metric direction before any model is run?
  • Are the same splits and preprocessing used for both models when comparing them?
  • Is the aggregate computed as a weighted average of per-subset scores, not as a single pooled metric?
  • Does the report include per-subset results so a high aggregate cannot hide a low score on a critical subset?
  • Are metric directions recorded, such as accuracy and recall maximizing versus error metrics minimizing?
  • Is the evaluation set versioned or described so results remain comparable across runs?
  • Does the comparison include a statistical check if the difference is small or noisy?
  • Is the final test set kept separate from the evaluation set used for model selection?
  • Are the scoring rules documented in a file or configuration that others can reuse?
  • Does the report show the raw per-subset numbers, not only the aggregated value?

This approach does not replace a separate final test set when hyperparameters are tuned. It does not by itself provide statistical significance; a formal test is needed for small or noisy differences. Results remain comparable only if the evaluation set and rules stay unchanged, because changing splits or weights requires a new comparison. It also does not remove the need to choose the validation strategy, such as cross-validation or chronological splits, based on the data type. The method assumes multiple subsets or metrics exist and that the decision is based on model comparison. It cannot guarantee that a weighted aggregate is the right business objective, and it does not align embedding spaces or establish relevance ground truth without paired encoders and independent reference labels.

Sources

  1. Hugging Face: evaluation ↗
  2. scikit-learn: model evaluation ↗
Back to top ↑