Trialling a receipt app on your own data

Nobody can tell you how well a receipt scanner will work on your receipts, including the people who built it. Performance depends on the vendors you buy from, the state of the paper by the time you photograph it, and whether the fields you need are the fields it extracts well. The only useful evaluation is a short trial on your own documents, scored deliberately.

Doing it properly takes an afternoon of setup and two weeks of ordinary use. Skipping it means choosing on the strength of a claim that cannot be checked.

Build the test set first

Assemble the receipts before you sign up for anything, so the same set goes through every tool you are comparing.

Include, deliberately, the awkward ones: a faded thermal slip, a long grocery receipt, one from a vendor whose name is nothing like their trading name, a handwritten total, one in a foreign currency, an email receipt, a crumpled one you found in a coat pocket. Add a few clean ordinary ones so the set is not pure worst case.

Then — and this is the step that makes the rest work — write down the correct answer for each one yourself, by hand, before any software sees it. Total, date, vendor, tax if you need it. Without this you will end up grading the software against its own output, which is how a trial concludes that everything is fine.

A set in the low tens is enough. Precision is not the goal; distinguishing “usable” from “not” is.

Score per field, and score three outcomes

A per-receipt pass rate hides everything you need to know. One tool may read totals flawlessly and dates badly, which is fine if you are reconciling by amount and fatal if you are filing by period.

Grade each field into three buckets, not two:

Right. Matches what you wrote down.

Wrong and flagged. Incorrect, but the tool marked it low-confidence, failed an arithmetic check, or otherwise sent it for review. This is a cost, not a failure — the system did its job.

Wrong and silent. Incorrect, high confidence, posted without comment. This is the only category that genuinely matters. A tool with more flagged errors and fewer silent ones is better than the reverse, and a single accuracy percentage cannot express that, which is why vendors quote one.

Count the silent errors per field. That number is your actual comparison.

Run the whole chain, not just the scan

Extraction is one of five stages, and the later ones are where trials most often turn out to have been misleading — see how the chain fits together.

Categorisation. Did it code things sensibly, and could you correct it in a way that stuck? A correction that does not become a rule is a correction you will make every month.

Reconciliation. Connect the card feed if you can, and see whether matching works on your real statement descriptors rather than on tidy ones.

Export and integration. Push something into your accounting system, or at least export it. This is where half the disappointments live — what an integration actually does varies enormously between tools that describe it with identical words.

Retrieval. Two weeks in, find a specific receipt from the first day using only what you would remember about it. If you cannot, nothing else about the tool matters.

Test the second week, not just the first

The first week measures cold performance. The second measures whether the system learns, and that is the more important number for a tool you will use for years.

Buy from the same vendor again and see whether the correction you made stuck. Send a duplicate through deliberately and see whether it is caught. Send an amount over whatever threshold you set and see whether it routes for review. Check whether the same supplier arriving under a different name gets recognised or quietly becomes a second vendor.

A tool that starts adequate and improves beats one that starts impressive and never changes.

Measure the friction, because it decides everything

The most accurate tool in your comparison will lose to a less accurate one if capture takes too long, and this is not a small effect — it is the difference between a system you use and a folder of unscanned paper.

Time the capture, at the till, one-handed, with a bag in the other hand. Count the taps from opening the app to a filed receipt. Note whether it works with no signal, because a queue that fails offline fails exactly when you are travelling.

Then count how many receipts from the trial period you actually captured versus how many purchases you made. A tool with an excellent read rate that you remember to use half the time is worse than the opposite.

What a trial cannot tell you

Be honest about the limits, and ask rather than test:

Whether you can leave. Export everything on the last day of the trial and look at what you got. Fields without images, or images without the field data, is a partial export dressed as a complete one — the reason to keep originals in an open format.

What happens in year three. Storage limits, plan changes, whether old receipts stay retrievable.

Long-run accuracy. Two weeks of your receipts is a small sample and a seasonal one.

What a pass looks like

Silent errors on the fields you rely on, near zero. Flagged errors at a rate your review time can absorb. Capture fast enough that you do it every time. Corrections that persist. An export you would be content to depend on.

If a tool clears those, the remaining differences are preference. If it fails the first one, no feature list compensates — and you will only know because you wrote the answers down before you started.