Choose a classification threshold by the cost of false positives and false negatives
A practical guide to selecting a decision threshold by weighting false‑positive and false‑negative costs, using a held‑out validation set and scikit‑learn confusion matrices.
On this page
The short answer
Compare candidate thresholds on validation data using explicit error costs. For the hypothetical counts in this guide, variant A has four false positives and one false negative; variant B has one and three. With costs 1 and 5, their totals are 9 and 16, so A is cheaper among these two options. This is an illustrative calculation, not a fitted model result or proof of the globally optimal threshold. Keep final test data separate and evaluate uncertainty and deployment conditions.
Separate scores from decisions
A classifier produces a score, while a threshold turns that score into a decision. For a binary rule, you can define the positive decision as score greater than or equal to the threshold. The direction of the score and the positive class must be explicit. A probability-looking value does not by itself establish that the model is calibrated.
Changing the threshold changes which cases receive positive and negative decisions. It does not retrain the underlying model or make its probabilities unbiased. A conventional cutoff such as 0.5 is a starting point for probability scores, not a universal cost-optimal setting. Compare alternatives against the decision you actually need to make.
Specify which error costs more
In many domains a false negative (FN) is far more damaging than a false positive (FP). Cost‑sensitive threshold selection requires the user to assign numeric penalties, e.g. cost(FP)=1 and cost(FN)=5. These numbers are not derived from the data; they reflect domain expertise, regulatory impact, or downstream expenses. Once the costs are fixed, the total cost for a given threshold is simply cost(FP)·FP + cost(FN)·FN.
Count the confusion matrix
Scikit‑learn's confusion_matrix follows the convention rows = true class, columns = predicted class. For binary problems with labels=[0,1], the entries are: TN = C[0,0], FP = C[0,1], FN = C[1,0], TP = C[1,1]. Using the validation scores and a candidate threshold we can obtain y_pred = (probs >= thr).astype(int) and then call confusion_matrix(y_val, y_pred, labels=[0,1]).
Compare two illustrative thresholds
Use explicitly hypothetical confusion counts: variant A has FP=4 and FN=1, and variant B has FP=1 and FN=3. Let a false positive cost one unit and a false negative cost five units. The total for A is 1*4 + 5*1 = 9. The total for B is 1*1 + 5*3 = 16.
A is cheaper among these two constructed options even though B has fewer false positives. These counts are supplied for arithmetic illustration; the code does not train a classifier or produce them from a dataset. The two printed totals below are expected results derived from the formula, not observations from an executed test.
cost_fp = 1
cost_fn = 5
fp_a, fn_a = 4, 1
fp_b, fn_b = 1, 3
cost_a = cost_fp * fp_a + cost_fn * fn_a
cost_b = cost_fp * fp_b + cost_fn * fn_b
print("Cost A:", cost_a)
print("Cost B:", cost_b)
Choose on held‑out validation
Train the model on training data, then compare thresholds on separate validation predictions. Do not calculate the validation counts using the same observations that fitted the model. Use an appropriate split for the task: a chronological problem may require a time-respecting split rather than random sampling.
Threshold selection itself uses the validation results, so reporting the smallest observed validation cost can be optimistic. A separate split prevents direct reuse of training examples for tuning; it does not guarantee an unbiased estimate or enough positive cases. Record the split, candidate thresholds, class counts and cost assumptions so the comparison can be interpreted.
Keep threshold selection away from the test set
Freeze the model, threshold and preprocessing before evaluating on final test data that were not used for fitting or threshold selection. If the test results cause you to tune again, that sample becomes part of the selection process and no longer serves the same independent evaluation role.
An independent test set is useful but not a guarantee of production performance. A small sample, an unrepresentative population or changing conditions can distort the estimate. Report the observed counts, sample size and uncertainty rather than calling a single total a guaranteed or unbiased production cost.
Check deployment prevalence
Deployment prevalence is the proportion of positive cases in the population where the decision will be used. It may differ from validation prevalence. To compare per-case expected costs under a proposed prevalence p, use cost_fp*(1-p)*FPR + cost_fn*p*FNR, where FPR is the fraction of negative cases classified positive and FNR is the fraction of positive cases classified negative.
This adjustment assumes that the conditional error rates remain applicable under the changed class distribution. It does not handle arbitrary changes in the features, calibration or labeling process. Both classes need sufficient observations to estimate their rates. Multiplying per-case expected cost by a proposed population size gives a total under these assumptions; do not compare raw counts from differently sized populations as if they were interchangeable.
State what the cost comparison omits
The simple cost formula ignores several real‑world factors: (1) the monetary value of downstream actions (e.g., additional tests), (2) the uncertainty in cost estimates, (3) potential benefits of early detection beyond binary outcomes, and (4) the effect of calibration - probability estimates may be biased, making a low threshold appear cheaper than it truly is. These omissions should be documented, and sensitivity analysis can be performed by varying the cost parameters.
Things to check
- Define the positive class and score direction explicitly.
- Keep training and threshold-selection observations separate.
- With labels=[0,1], read the confusion matrix as TN, FP, FN, TP in row-major order.
- State the false-positive and false-negative costs, including their units.
- Treat the 9 and 16 totals as constructed arithmetic examples.
- Keep final test data out of threshold selection.
- When adjusting prevalence, state the assumption about conditional error rates.
Where this applies
The example uses hypothetical binary confusion counts and fixed costs. It identifies the cheaper of two options, not a globally optimal threshold. Independent validation or test splits do not guarantee unbiased production estimates, calibrated probabilities or stability under distribution changes. Real deployments need representative data, uncertainty analysis and justified cost assumptions.