doc2data test evidence
This page publishes the results behind every doc2data number on this site, so you can check them. All test documents are synthetic: generated by our own script, with invented companies, addresses and VAT numbers. These are not results on real supplier invoices, and nothing here comes from a client.
Automated tests
21 of 21 passed during the build.
Results by set (synthetic)
| Set | How it was built | Documents | Fields correct | Field errors | Correctly flagged (deliberate error) | False alarms | Silent errors |
|---|---|---|---|---|---|---|---|
| Main | Development set: the rules were written against these layouts | 16 | 189 of 189 | 0 | 1 of 1 | 0 of 15 | 0 |
| Held-out values | Same layouts, new values | 16 | 191 of 191 | 0 | 1 of 1 | 0 of 15 | 0 |
| Third set | New values; per the developer, measured before (174 of 178) and after the OCR change and not deliberately used to tune it; included in automated tests throughout development | 16 | 178 of 178 | 0 | 1 of 1 | 0 of 15 | 0 |
| Degraded scans | Every invoice as a degraded scan; used while developing a tilt (deskew) fix | 12 | 131 of 141 (92.9%) | 10 | 1 of 1 | 6 of 11 | 1 |
| Degraded scans, hold-out | Same damage, new values; per the developer's log, not looked at during that fix | 12 | 127 of 139 (91.4%) | 12 | 1 of 1 | 7 of 11 | 0 |
How to read the columns
- Fields correct: extracted fields that match the known answers that the generator created: 8 fields per document (vendor, VAT number, invoice number, date, currency, net, VAT, total) plus every line item. For example, main = 8 × 16 + 61 line items = 189.
- Correctly flagged: each set contains one invoice with a deliberate arithmetic error, and it went to the review list every time.
- False alarms: otherwise valid documents that were also put on the review list, mostly because a total could not be read.
- Silent errors: documents that passed every check but contained a wrong field. The one case was a vendor name read as "S.n.C." instead of "S.n.c.".
Field by field on the degraded sets (synthetic)
| Field | Degraded scans | Degraded scans, hold-out |
|---|---|---|
| Vendor | 10 of 12 | 11 of 12 |
| Vendor VAT number | 11 of 12 | 12 of 12 |
| Invoice number | 12 of 12 | 12 of 12 |
| Date | 12 of 12 | 12 of 12 |
| Currency | 12 of 12 | 12 of 12 |
| Net | 12 of 12 | 11 of 12 |
| VAT | 12 of 12 | 12 of 12 |
| Total | 8 of 12 | 7 of 12 |
| Line items (description, quantity and amount match) | 42 of 45 | 38 of 43 |
How the sets were built
- A generator script creates invoices and receipts with known correct values. Each set uses a different random seed (7, 99, 123, 2024 and 5150).
- Clean sets (16 documents each): 12 invoices in Italian, English, German and French with 4 currencies (EUR, GBP, CHF, USD), 2 receipts and 2 clean scans (image copies of 2 of the invoices). They come from 2 invoice layouts plus a receipt format.
- Degraded sets (12 documents each): every invoice is rendered as a scan at 110 dpi, tilted by 1.5°, blurred and saved as a JPEG at quality 30.
- One invoice per set contains a deliberate arithmetic error that should be flagged.
Limits of this evidence
- No unseen layouts. All sets share the same generated layout families. The held-out sets change values, not layouts.
- Small samples. With 16 or 12 documents per set, zero errors on the clean sets supports no accuracy bound for real invoices. For the two degraded scores (92.9% and 91.4%): these small synthetic sets do not establish a reliable difference in general performance; no analysis accounting for document clustering has been performed. Fields within a document are not independent.
- The hold-out record comes from the developer's log. No dated record shows it was frozen in advance. The third set was part of the automated tests (pass mark 95%) throughout development.
- Platform differences. On Windows, the developer measured 95.7% and 94.2% on the two degraded sets. The Docker (Linux) figures above are the reference. This is a single-environment observation, not a platform comparison.
- Not tested: real supplier documents, phone photos, handwriting, multi-page tables, several VAT rates on one invoice, credit notes, and the optional AI fallback against a live model.
- Real-world accuracy has not been measured and may differ, including being lower.
Check it on your documents
The only accuracy that matters for you is on your own documents. Send 3 samples for a free feasibility check: a first look, not a reliable estimate. Send 3 sample documents Back to doc2data
How the optional AI fallback works (code reviewed 3 October 2026)
- Off by default.
- Cloud option (command-line tool only): asks for the fields the rules missed, and its prompt includes up to the first 6,000 characters of the extracted text. That text goes to the OpenAI-compatible provider you pick, using your own API key.
- Local option (Ollama): The optional Ollama fallback runs on the same machine (localhost); nothing leaves the computer.
- The demo app offers only the local option.
- Both options are built and unit-tested with a simulated backend; neither has been tested against a live model.
Source: transcription of Snello's private developer reports (accuracy.json for each set) and code review, not an independently reproducible public benchmark. Reports generated 2 October 2026 at commit e4bfc58; re-checked by company-engineer the same day by rebuilding the Docker image at commit 708168d (identical results); currency description corrected in commit 2eddaff. The 174 of 178 pre-change result and the hold-out history are as reported by the developer. Figures transcribed for this page on 3 October 2026.
Bring your workflow into focus.
Send up to 3 samples or describe one process. We reply by email with what is feasible and how we would measure it.