Synthetic transaction data
How the invented bank-statement descriptors are generated, and why the train, validation and test splits are made by merchant rather than by row.
What it does
A seeded generator invents 24,000 bank-statement lines in 14 spending categories. Each line has a descriptor (the text a bank prints, such as "SQ *BLUE BOTTLE COFFEE SAN FRANCISCO CA"), a date, and usually an amount. Merchants are either well-known public brands or local businesses assembled from generic word lists. The people in peer-to-peer transfers and the employers in payroll deposits are made up.
The generator imitates how card and ACH descriptors look: processor prefixes ("SQ *", "TST*", "PAYPAL *"), store numbers, a city and state, reference numbers, truncation to a fixed field width, abbreviations, dropped punctuation and mixed casing. Amounts follow a distribution per category, and one row in ten has no amount at all.
Why it is used
Real transaction data is personal financial data. Publishing it, or a model trained on it, is not an option, and no public dataset of labeled statement descriptors is large and clean enough to use. Synthetic data makes the whole pipeline reproducible and shareable: anyone can regenerate the exact dataset from the seed.
The split is the part that matters most. Every transaction from one merchant, and from its sister brands (for example a ride-hailing app and its food-delivery arm), goes to exactly one of training, validation or test. The test set therefore measures what the model does with a merchant it has never seen, which is the only hard case in practice: a merchant already seen can be looked up.
Inputs
- A fixed seed (20260928), so the dataset is identical on every machine.
- About 210 public brand names, each with a category and descriptor templates.
- Generic word lists for local businesses (for example "family dental", "taqueria", "auto repair"), about 70 invented merchants per category.
- Split shares of 70% training, 15% validation and 15% test, assigned by merchant group.
Assumptions
- The descriptor noise (prefixes, truncation, casing, store numbers) resembles what US banks print. It was written from public examples of statement formats, not measured on real statements.
- Category frequencies and amount ranges are plausible but invented.
- Each merchant has one true category, except a few deliberately ambiguous chains (a supercenter receipt can be groceries or shopping) whose label is drawn at random per transaction. That puts a floor under the error rate, as with real data.
How to read the results
The table below counts rows, merchants and merchant groups in each split. "Merchant groups shared across splits" must be zero: that is the leakage check. Test counts per category show how much evidence stands behind each per-category score; categories with a few hundred test rows from a handful of merchants have wide uncertainty.
Limitations
- Synthetic text is cleaner and more regular than real statements. Scores here are an upper bound on what the same model would reach on real bank data.
- The brand list is finite. A real deployment meets thousands of merchants the generator never imagined.
- Only US-style descriptors in English.
Where it can fail
- If the generator's templates are too distinctive per category, the model learns the templates rather than the language of merchants. The by-merchant-type results on the limitations page are the check: local businesses built from category words are easy, unseen brand names are not.
Validation on current data
Counts from the dataset the running engine's model was trained and tested on.
References
- Kaufman, S., Rosset, S., Perlich, C. and Stitelman, O. (2012). Leakage in data mining: formulation, detection, and avoidance. ACM Transactions on Knowledge Discovery from Data 6(4).
- Jordon, J. et al. (2022). Synthetic data: what, why and how? The Royal Society and The Alan Turing Institute.