Reading Arabic documents field by field
multi-sector enterprise · Gulf
The shape of this problem
A human reads it and types it in again
This is one of our own builds, not a client engagement. It is capability evidence and it is described as such.
The problem
Organisations processing Arabic documents needed reusable capabilities for entity extraction, summarisation, classification, search, and structured information retrieval. In practice every Arabic document project rebuilt the same foundations - normalisation, OCR handling, label sets, evaluation - before reaching anything specific to the customer. The cost of finding out whether a use case was viable therefore approached the cost of delivering it, which is what kept most of these projects from starting.
What we built
We built an accelerator combining Arabic named-entity recognition, document classification, summarisation, OCR-ready preprocessing, semantic embeddings, and configurable output schemas, with notebook experiments and packaged components for rapid evaluation against customer samples before production integration.
What moved
| Measure | Before | After |
|---|---|---|
| Manual document-review time | † | down 60% |
| Classification consistency | † | up 30% |
| Structured-field extraction coverage | † | up 28% |
† Measured against the prior level of the same measure. The source publishes the size of the movement, not the figure it moved from.
Manual document-review time
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Classification consistency
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Structured-field extraction coverage
Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.
Figures are drawn from the practice's own delivery records for the engagement named, measured against the process that preceded it. They have not been through third-party audit, and none is presented as an average across clients.
What it turned on
Coverage rose because OCR, entity recognition, page zones, and validation rules were combined rather than expecting one model to recover every value - and page-level provenance plus review routing kept higher coverage from meaning that more weak predictions were silently accepted as authoritative business data.
- Service line
- Bilingual Document AI
Start here
Which of the four is yours?
Tell us the documents and the monthly volume and we will send the two closest records, with the proof behind each and the honest note where the match is partial.
You get a reply within one working day, from the engineer who would do the work - not a sales sequence.