What is Document extraction accuracy?
A single accuracy figure for document extraction is close to meaningless. The split by input quality is the real number.
Extraction accuracy varies enormously by what arrives. On a delivered system reading scanned paper, the same pipeline ran at 96% on clean digital forms, 84% on poor scans and 61% on handwritten entries. Averaging those produces a number in the eighties that describes none of the three, and which is wrong in the direction that flatters the vendor.
The question that actually matters is what happens to the uncertain remainder. On that same system, 14% of fields were routed to human review, each shown alongside the region of the page it came from so a reviewer could confirm it in seconds. A vendor who quotes one figure and cannot describe the review path has either not measured it across real inputs or has chosen not to say.
| On this | Document extraction accuracy | A single quoted accuracy figure |
|---|---|---|
| What it tells you | How the system behaves on each kind of document you actually have | A weighted average of a document mix that is probably not yours |
| What it hides | Nothing - the weak case is stated | The handwritten and poor-scan cases, which are usually the expensive ones |
| The remainder | Routed to review with its page location, so it costs seconds | Unaddressed, which means somebody re-reads the whole document |
- When it is the right answer
- Ask for the split before you ask for the figure. If it cannot be produced, the figure has not been measured on real inputs.
- The service line that delivers this
- AI Document Intelligence
Four thousand supplier invoices a month, and three people typing them into the ERP. Month-end waits for them.
Start here
If a term here is the one your board paper turns on, ask us about it.
We will send the entry, the evidence behind it, and the honest note about where it does not apply.
You get a reply within one working day, from the engineer who would do the work - not a sales sequence.