Features and model
Character n-grams, words and amount flags feeding a temperature-scaled multinomial logistic regression, with an exact per-word explanation of every prediction.
What it does
Each descriptor is normalized (lower case, accents folded, punctuation removed, every digit mapped to 0 so store and reference numbers keep their shape but cannot be memorized) and turned into three blocks of features: character n-grams of 3 to 5 characters within word boundaries, words and word pairs, and a handful of amount flags (debit or credit, a size bin, whole dollars, a .99 or .95 ending, or "amount missing"). The text blocks are TF-IDF weighted with sublinear term frequency and scaled to unit length per block.
A multinomial logistic regression turns the features into one score per category, and a softmax turns the scores into probabilities. The probabilities are divided by a temperature T fitted on validation merchants (temperature scaling), which fixes over- or under-confidence without changing which category wins.
Why it is used
Character n-grams survive the damage statement formats do to names: "STARBUCKS #1234" and a truncated "STARBUCK" share most of their n-grams, and so do "COFFEE" and "COFFE". Words add meaning that n-grams blur, such as "family dental".
A linear model is used because it is explainable exactly. A category's score is a sum over features, so every prediction can be broken down into how much each word and the amount pushed toward or away from each category, and the pieces add up to the score. On short texts like these, linear models on n-grams are also hard to beat by much.
Inputs
- A descriptor of up to 200 characters.
- An optional amount in dollars (negative for money going out).
- The trained artifact: vocabulary, IDF weights, coefficients (stored as float16), intercepts and the temperature.
Formulas
Assumptions
- Word order beyond pairs does not matter.
- Features add up: the model cannot learn that a word means one thing next to another word and something else alone, beyond what word pairs capture.
- The regularization strength C is chosen on validation merchants by macro-F1, and the temperature by log loss on the same merchants. The test set is used once, at the end.
- Training runs offline with scikit-learn (L2-penalized, L-BFGS). The site runs inference in NumPy from the saved weights, so scikit-learn is not needed to serve predictions.
How to read the results
The confidence shown is the calibrated probability of the top category. The explanation colors each word by how much it moved the score of the chosen category: toward it, or away from it. Contributions are measured against the average across categories, because adding the same number to every category's score changes nothing.
A low share of known features (text the model has never seen) is a warning sign: the model is then guessing mostly from the amount and generic fragments.
Limitations
- An explanation shows what the model used, not why a category is right. A confident prediction built on a store-number fragment is still a guess.
- Weights are stored as float16 to keep the artifact under 1 MB. The published test metrics are recomputed from the stored float16 weights, so they describe the model that is actually served.
- The vocabulary is fixed at training time. New words carry no weight until the model is retrained.
Where it can fail
- A brand name made of another category's words (say, a clothing brand called "Harvest Kitchen") is classified by its words, confidently and wrongly.
- Very short descriptors ("PAYMENT", "POS 0412") carry almost no signal; the prediction then leans on the amount.
Validation on current data
The regularization search on validation merchants, and ablations measured on the test merchants: each feature block alone and in combination.
References
- Jurafsky, D. and Martin, J. H. Speech and Language Processing, 3rd ed. draft (chapters on logistic regression and naive Bayes text classification).
- Guo, C., Pleiss, G., Sun, Y. and Weinberger, K. Q. (2017). On calibration of modern neural networks. ICML.
- Pedregosa, F. et al. (2011). Scikit-learn: machine learning in Python. Journal of Machine Learning Research 12.