Modern models read invoices and quotes well. What they can't do is carry the consequences of being wrong. So every value that matters still gets checked by a person, and that checking now sets the real cost of intelligent document processing. This briefing sets out the measured evidence, and what it changes about how document AI is bought, piloted and reviewed.
Four of five extraction approaches returned flawlessly formed output on every document. Structure told us nothing about truth.
Four approaches returned flawless structure on every document. Their field accuracy still differed by 11 points. The metric everyone checks could not tell them apart.
The strongest of five approaches still fell short of ground truth on roughly one field in eleven, with no indication of which ones. Human review is till essential.
These conclusions come from measurement, not positioning. They shaped the evaluation, and they should shape the buying conversation.
A well-formed record can still contain inaccurate data. The wrong value may pass validation, load without errors, and go unnoticed, until it affects a payment, a certificate, or an audit.
Each question is phrased the way a leader would actually ask it. A system that answers yes to all four is audit-ready. A system that answers with an accuracy percentage is answering a different question.
We define the critical fields, who reviews them, and what a reliable answer must include before selecting any tool.