Evaluation and leakage control

Accuracy, macro-F1 and calibration on unseen merchants, compared with keyword rules and a majority guess, with merchant-level bootstrap intervals and a measured leaky split.

What it does

The model is scored once on the test merchants. Accuracy is the share of rows categorized correctly. Macro-F1 averages the F1 score (the harmonic mean of precision and recall) over the 14 categories, so a small category counts as much as a large one. Calibration is measured with the expected calibration error (ECE): predictions are grouped into ten confidence bins, and ECE is the row-weighted average gap between the confidence and the accuracy in each bin.

Two baselines put the numbers in context: always guessing the most common category, and a list of keyword rules of the kind a person would write ("UBER" is transport, "PAYROLL" is income). A third line uses the rules when one matches and the model otherwise.

Why it is used

A model only earns its complexity if it beats the simple alternatives on the hard case. Keyword rules are what most budgeting tools start with, so they are the baseline that matters.

The same model is also scored on a random row split, where most test merchants were seen in training. The gap between the two numbers is the size of the mistake the merchant split avoids.

Inputs

  • Test rows: every transaction from 15% of merchant groups, none seen in training or validation.
  • The model's calibrated probabilities and the baselines' predictions for the same rows.
  • 1,000 bootstrap resamples of merchant groups for the intervals.

Formulas

Precision_k = TP_k / (TP_k + FP_k) Recall_k = TP_k / (TP_k + FN_k) F1_k = 2 · Precision_k · Recall_k / (Precision_k + Recall_k) Macro-F1 = mean_k F1_k ECE = Σ_b (n_b / n) · | accuracy_b − confidence_b | (10 equal-width bins) Log loss = −(1/n) Σ_i ln P(true category of row i)

Assumptions

  • Rows from the same merchant are not independent: a model that misses a brand misses all of its rows. The bootstrap therefore resamples whole merchants, which gives wider and more honest intervals than resampling rows.
  • The gain over the keyword rules uses the same resamples for both, so it is a paired interval.
  • Temperature scaling is fitted on validation merchants and never sees test rows.

How to read the results

Read the interval before the point estimate. With about 170 test merchants, one brand family can move accuracy by a few points, and the intervals show it.

ECE near zero means the confidence can be taken at face value: of predictions made at 80% confidence, about 80% are right. The coverage table on the limitations page shows the practical use: accept confident predictions automatically and send the rest to a person.

Limitations

  • All numbers come from synthetic data. They show the method working and failing in controlled conditions, not what it would score on a real bank's data.
  • The keyword rules were written by the same person who wrote the generator, so they are a fair but not an adversarial baseline.
  • Rules-first can beat the model alone here because the rules were written for the very brands in the generator's list. On unseen merchants rules only help when a category word appears in the name.

Where it can fail

  • A random row split would report near-perfect accuracy for a model that has only memorized merchant names. The table below shows by how much.
  • Macro-F1 on a category with few test merchants rests on very little evidence.

Validation on current data

Test-set results for the model the running engine serves. The test suite recomputes them from the stored weights and checks the metric code against scikit-learn's.

References

  • Naeini, M. P., Cooper, G. F. and Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. AAAI.
  • Efron, B. and Tibshirani, R. J. (1993). An Introduction to the Bootstrap. Chapman and Hall.
  • Kapoor, S. and Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4(9).

Try it in the categorizer