Reading Arabic documents field by field
multi-sector enterprise · Gulf
The shape of this problem
A human reads it and types it in again
Built and run by this practice in its own environment.
The problem
Organisations processing Arabic documents needed reusable capabilities for entity extraction, summarisation, classification, search, and structured information retrieval. In practice every Arabic document project rebuilt the same foundations - normalisation, OCR handling, label sets, evaluation - before reaching anything specific to the customer. The cost of finding out whether a use case was viable therefore approached the cost of delivering it, which is what kept most of these projects from starting.
What we built
We built an accelerator combining Arabic named-entity recognition, document classification, summarisation, OCR-ready preprocessing, semantic embeddings, and configurable output schemas, with notebook experiments and packaged components for rapid evaluation against customer samples before production integration.
What moved
| Measure | Before | After |
|---|---|---|
| Manual document-review time | † | down 60% |
| Classification consistency | † | up 30% |
| Structured-field extraction coverage | † | up 28% |
† Measured against the prior level of the same measure.
Manual document-review time
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Classification consistency
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Structured-field extraction coverage
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Figures are from the practice's own delivery records for the engagement named, measured against the process that preceded it.
What it turned on
Coverage rose because OCR, entity recognition, page zones, and validation rules were combined rather than expecting one model to recover every value - and page-level provenance plus review routing kept higher coverage from meaning that more weak predictions were silently accepted as authoritative business data.
- Service line
- Bilingual Document AI
Start here
Which of the four is yours?
Tell us the documents and the monthly volume and we will send the two closest records, with the figures behind each and what it would take to repeat them on your material.
You get a reply within one working day, from the engineer who would do the work.