Insights

Data, analytics & ML evaluation

One lucky split: how to tell if your model is really better

An illustrative case: your new model scores 0.87 on the test set, and the old one scored 0.85. Should you switch? Maybe. But the 0.02 gap might depend on which rows happened to land in the test set. Split the data differently and the ranking could change. Here is a more careful way to look at it before a decision depends on it.

Published Snello (Foysal)Drafted with AI tools; Foysal is responsible for the content.Method from our open-source project Bayesian-Classifier-Lab. All numbers in the examples below are illustrative, not results of a named experiment.

Why one split is not enough

A single train/test split gives one number per model. That number mixes the model's skill with the luck of the split: unusually easy test examples, or a rare class that ended up mostly in training. With one split you cannot separate the two, and with small or imbalanced datasets the luck can be large.

Step 1: repeat the experiment, with a suitable design

Cross-validation splits the data into k parts (folds), trains on k−1 and tests on the remaining one, then rotates. Repeated cross-validation does this several times with different shuffles, so each model gets many scores instead of one.

Some details matter:

  • Use the same folds for every model. Shared folds align the evaluation data for a paired comparison. How much a hard fold affects each model can still differ.
  • Keep every fold score, not just the average. The spread describes how variable the scores were; it is not by itself a confidence interval for future performance.
  • Choose the split design to fit the data. Ordinary shuffled splits assume rows can be treated as exchangeable. If one customer or company appears in many rows, use group-aware splits. If the data are ordered in time, use time-aware splits. Stratified folds keep class proportions similar across folds; they do not fix every imbalance problem or make rows independent.

Libraries such as scikit-learn provide stratified, grouped, time-series and repeated splitters.

Step 2: ask the question you care about

The usual question is "is the difference statistically significant?" It has two problems. With enough data, tiny and useless differences become "significant". And a non-significant result is often misread as "the models are the same".

A more useful question is: how probable is it, under a stated model, that A is better by a margin that matters to us?

Step 3: a Bayesian comparison with a "doesn't matter" zone

Bayesian-Classifier-Lab uses a Bayesian correlated t-test with a region of practical equivalence (ROPE), following Benavoli and colleagues (2017):

  • Decide first what difference is too small to matter, for example ±0.01 in accuracy (illustrative). That band is the ROPE. Choosing it is a business decision.
  • Compute the paired differences between the two models on every fold.
  • Account for overlap, approximately. Cross-validation folds share training data, so their scores are correlated. The correlated t-test uses a heuristic correlation of ρ = test-set size / total size, related to the correction of Nadeau and Bengio (2003), together with a flat (non-informative) prior. This reduces the overconfidence of treating folds as independent; it does not remove all dependence or guarantee well-calibrated uncertainty.
  • Read three probabilities: that A is better by more than the ROPE, that the two are practically equivalent, and that B is better by more than the ROPE. They are conditional on the model and prior above. They are not measured probabilities of business success.

Then decide with the assumptions in view. Illustrative examples:

  • "96% probability that A is practically better" would meet the decision threshold of 0.95 that Bayesian-Classifier-Lab uses by default. Whether to switch still depends on the cost and risk of switching.
  • "60% equivalent, 30% A, 10% B" suggests the new model is probably not worth the change.
  • "Inconclusive" is an allowed answer. Before collecting more data, check for leakage, grouping or time-order problems in the split design.

If you set a different threshold, state it and why.

Make it reproducible

Bayesian-Classifier-Lab's documentation and code describe these practices:

  • fixed random seeds, with split settings in a configuration file;
  • each fold's identifier, score and timing saved, so figures can be traced back;
  • a run record with the software versions used;
  • preprocessing (for example scaling) fitted inside each training fold, so test data do not leak into training;
  • comparisons rejected when fold keys don't match, rather than quietly comparing different splits.

These are documented capabilities of the project. They are not a new, independent validation of it.

Common traps

  • Tuning on the reported folds. If hyperparameters were chosen on the same folds you report, scores are optimistic. Use nested cross-validation or a separate validation split.
  • Comparing many models and reporting only the winner. With enough candidates, one looks best by chance. Report every comparison.
  • Ignoring calibration. Two models with equal accuracy can give very different probability estimates. If decisions use the probabilities, check calibration too.

Try it yourself

The framework is open source: Bayesian-Classifier-Lab on GitHub. It compares five classifier families on clean tabular data and includes a synthetic demo dataset, which shows the workflow and is not evidence about any real-world task.

The takeaway

Before trusting "model B beats model A":

  1. Repeat the split with a design that fits your data, using the same folds for every model.
  2. Decide in advance how big a difference matters, and what probability you require.
  3. Report probabilities with their assumptions, including "practically equivalent" and "inconclusive".

Need a second opinion on a model? We offer this evaluation, with calibration checks and explanations, as part of our data and ML evaluation service. Free feasibility check

Sources

  • Benavoli, A., Corani, G., Demšar, J., Zaffalon, M. (2017). "Time for a Change: a Tutorial for Comparing Multiple Classifiers Through Bayesian Analysis." Journal of Machine Learning Research 18(77):1–36. https://jmlr.org/papers/v18/16-305.html (retrieved 3 October 2026).
  • Nadeau, C., Bengio, Y. (2003). "Inference for the Generalization Error." Machine Learning 52:239–281. https://link.springer.com/article/10.1023/A:1024068626366 (retrieved 3 October 2026; abstract and metadata).
  • scikit-learn user guide, "Cross-validation: evaluating estimator performance", https://scikit-learn.org/stable/modules/cross_validation.html (retrieved 3 October 2026).
  • Bayesian-Classifier-Lab README and source (decision threshold default 0.95; correlation ρ = n_test/n_total), https://github.com/Foysal-A-Al/Bayesian-Classifier-Lab (retrieved 2–3 October 2026, as reviewed by company-senior-researcher).

Bring your workflow into focus.

Send up to 3 samples or describe one process. We reply by email with what is feasible and how we would measure it.

Start a project
SNELLO / TOOLS

This explains the tool; it is not a live AI session. No document is uploaded.

Our working method
Full diagram

Read the project

Discuss your project

Send up to 3 samples or describe one process, the tools you use and the result you need.

snello.contact@gmail.com

Go to the Contact page

Navigation