100% Accurate Invoice Template Detection: A DoxTract Benchmark Case Study
Most OCR and IDP vendors quote a single accuracy number and leave it there. We wanted to know what actually happens field by field, template by template, when SoceTonAI DoxTract processes real invoices at scale — so we ran a benchmark and published the full methodology and results.
Here's what we tested, what we found, and what it means if you're evaluating DoxTract for invoice or accounts payable automation.
The Setup
We built and benchmarked DoxTract — a template-driven document extraction system — on a real-world public dataset of 7,000 invoice images from Kaggle, using a subset of 4,000 invoices across 4 distinct templates to simulate real enterprise variability in layout and vendor format.
For each template, we evaluated three fields that matter most for accounts payable workflows:
Invoice Number
Total Amount
Contact / Vendor
Each template was tested independently, using the same freemium extraction workflow available to any DoxTract user — draw the template once, save it to the cloud, and extract at scale.
What We Found
On stable, consistent layouts, extraction was close to perfect. Template 7 in particular showed near-deterministic performance — once a template is correctly mapped to a layout that doesn't shift, DoxTract extracted fields with up to 100% accuracy and minimal post-processing.
Invoice Number and Contact/Vendor held up strongly across every template, ranging from 92% to 100% accuracy for Invoice Number and 94% to 100% for Contact — both are structurally stable fields (a fixed label, a predictable position near the top of the page), and the results reflect that.
Total Amount was the clear outlier — and the most instructive result. Across layout-shifted templates, Total Amount accuracy dropped as low as 67%. Template 4 illustrates this most clearly: Contact extraction landed near-perfect at ~99.9%, Invoice Number at ~92.8%, but Total Amount fell to just 67.4%.
The cause wasn't OCR misreading characters — it was structural. Multi-line totals, tax and subtotal groupings, and positional variation in where the total actually sits on the page all confuse a fixed-template approach in ways that have nothing to do with text recognition quality. In other words: OCR wasn't the bottleneck. Layout understanding was.
Why This Result Matters More Than a Single Accuracy Number
It would have been easy to publish "DoxTract achieves 100% accuracy" and stop there — it's true, on the right templates. But that framing hides the more useful information: which fields are reliable by default, and which ones need a closer look before you trust straight-through processing.
That distinction is the entire point of running a field-level benchmark instead of a single blended accuracy score. A vendor quoting one number can't tell you whether your documents will land in the 100% bucket or the 67% bucket — because it depends entirely on how consistent your specific invoice layouts are, not on the OCR engine underneath.
The Cost Side of the Benchmark
The other half of the story is pricing. DoxTract's template-based extraction — no model training, no labeled dataset, no fine-tuning per vendor — runs as low as $0.6 per 1,000 pages, which includes OCR processing, template-based extraction, and structured JSON output. That's the direct benefit of defining structure once and reusing it indefinitely, rather than paying per-token inference costs on every document.
How It Works
The benchmark used DoxTract's standard three-step workflow, no custom engineering required:
Draw a template — visually map fields by connecting regions directly on a sample document (Invoice Number, Vendor Name, Total Amount, etc.). No training data needed.
Save the template to the cloud — store it for reuse across every future invoice that shares that layout, across teams and workflows.
Extract at scale — upload new invoices matching a saved template and get structured output instantly.
This is the practical difference between a template-reuse model and a model-training approach: once a template is defined, it doesn't degrade or need retraining — it either matches the new document's layout or it doesn't, and the benchmark results tell you which fields to trust when it does.
Where We're Focused Next
The benchmark surfaced a clear roadmap, not just a result. We're actively working on:
Improving extraction robustness for semi-structured invoices, where layout consistency is weaker than a strict template
Handling multi-layout vendor formats without requiring a brand-new template per minor variation
Closing the gap specifically on total-field extraction, since that's the field with the most room to improve
Try It on Your Own Documents
Benchmark numbers on a Kaggle dataset are useful, but the only number that matters for your team is what happens on your own invoices. DoxTract's free tier includes enough pages and templates to test that directly before committing to anything.
