A point of view on document AI

Extraction isn't the hard problem anymore. Verification is.

Modern models read invoices and quotes well. What they can't do is carry the consequences of being wrong. So every value that matters still gets checked by a person, and that checking now sets the real cost of intelligent document processing. This briefing sets out the measured evidence, and what it changes about how document AI is bought, piloted and reviewed. 

The numbers that should be on every operations agenda

Accuracy isn't the story.
The spread is.

1.000
Perfect structure, four times over

Four of five extraction approaches returned flawlessly formed output on every document. Structure told us nothing about truth. 

Read the analysis: Valid JSON is not correct data
0.798 - 0908
The spread behind the perfect structure

Four approaches returned flawless structure on every document. Their field accuracy still differed by 11 points. The metric everyone checks could not tell them apart. 

Read the analysis: Document AI has a new bottlenck
0.908
The best score anyone reached

The strongest of five approaches still fell short of ground truth on roughly one field in eleven, with no indication of which ones. Human review is till essential.

Read the analysis: The citation is the deliverable
What we believe

Three convictions shaping how we think about document AI in 2026.

These conclusions come from measurement, not positioning. They shaped the evaluation, and they should shape the buying conversation.

Conviction I

Verification outranks accuracy.

The model doesn't bear the consequences of its errors; your organisation does. So every value that matters gets checked by a person, at any accuracy level current tools reach. The binding cost is checking, not the reading. 
Conviction II

Valid is not correct

A well-formed record can still contain inaccurate data. The wrong value may pass validation, load without errors, and go unnoticed, until it affects a payment, a certificate, or an audit.

Conviction III

The citation is the deliverable.

An extracted value is not complete without visual evidence. In certification, the audit trail is part of the deliverable; in finance, it supports defensible decisions. Reviewers should verify the result, not search for its source.
Before the shortlist

Some options aren't ranked out. They're ruled out.

Privacy and capability tier are decided by your case, not by a vendor. Both narrow field before accuracy is worth discussing.
  • 1"Does this use case actually need a frontier reasoning model?"Model capability and extraction quality are not the same purchase. A clean, consistently structured document mix may not need frontier reasoning at all; a mix full of poorly structured tables may need more than a model on its own. The tier you choose is a cost that recurs per document, indefinitely. Establish what the work requires before assuming the top of the range, and test the assumption on your own documents rather than a vendor's.

Answer both and the shortlist shortens. What remains still has to prove it was designed for audit.
  • 2"Does the data need to stay on our infrastructure?"Some approaches send each document to a hosted service; others parse on hardware you control. If the data cannot leave, that is a constraint rather than a preference — it removes options instead of ranking them, and it does so before any accuracy conversation is meaningful. Processing locally has its own price in hardware and in time per document. Decide the boundary first, then evaluate what is left.
The four questions

Four questions. One honest answer.

Each question is phrased the way a leader would actually ask it. A system that answers yes to all four is audit-ready. A system that answers with an accuracy percentage is answering a different question.

1
Traceability
"Where did this number come from?"
Every extracted value should carry a citation to its exact source: the page and passage it came from. Confirming a cited value is a glance. Hunting for an uncited one is a search repeated per field.
2
Discrepancies
"What happens when the document disagrees with itself?"
Failed arithmetic has two distinguishable causes: the system misread it, or the document is genuinely wrong. The system must surface the discrepancy, never silently fix it. In regulated work, the discrepancy is the finding.
3
Attribution
"Can a failure be traced to a pipeline stage?"
Errors can occur during conversion, schema interpretation, or model reasoning. When a defect appears, its source should be clear, so teams can fix the problem instead of debating who is responsible.
4
The human gate
"Is there a checkpoint between extraction and business judgement?"
Verification should happen before action. A clear review step separates the two, helping teams catch extraction errors before they are mistaken for genuine business findings.
Our Method

We don't score demos. We score against ground truth.

What doesn't count
Demo impressions
and borrowed benchmarks
  • Clean-looking output on a handful of hand-pikced documents
  • Accuracy percentages measured on someone else's corpus
  • Schema-valid records read as correct records
What does
Answers known by construction
authored by ground truth
  • Documents authored as structured data first, with totals, quantities and references validated before rendering
  • Every field scored against ground truth that no tool under test produced
  • Every failure attributed to the pipeline stage that caused it, across five diagnostic views
How it works

We don't start with a tool. We start with what the documents are for.

1

Requirements

We define the critical fields, who reviews them, and what a reliable answer must include before selecting any tool.

2

Method selection

Candidates are shortlisted against those requirements — capability tier, data boundary, and whether citations are available at all. Direct-model extraction and parser-based pipelines are both on the table.
3

Testing

We measure performance, not demos. Using 39 synthetic invoices and quotes in both digital and image-only PDFs, we test how the tools really perform, trace every failure to its source, and keep confidential data protected.
4

Solution for the case

The approach that survives testing is then built around your documents and your review queue, with evidence tracing carried through the reviewer.
Why this matters now

Reviewers who search are the bottleneck.
Reviewers who aren't.

Without traceability
Searching
the reviewer sets the pace
  • The reviewer opens every document and hunts for the evidence behind every value
  • Silent errors are absorbed at whatever rate light sampling misses
  • Review headcount becomes the cost line that never shrinks
  • When something breaks, the diagnosis defaults to "the integration is broken"
With traceability
Confirming
citation sets the pace
  • Each value lands the reviewer on the exact span it came from
  • Documents that disagree with themselves arrive flagged, with both sides cited
  • The same reviewers clear more documents, and the queue that sets deployment cost gets shorter
  • Failures carry the address of the stage that caused them
Subscribe

Ideas like this shape how we work at Aicadium

Get updates about our latest AI Transformation projects.