Industrial AI practice  ·  Gulf  ·  Arabic and English Reply within one working day
Generative AI / retrieval-augmented generation

Answering a question across bilingual documents

multi-sector enterprise · Gulf

← All documented proof

The shape of this problem

The answer exists and nobody can find it

This is one of our own builds, not a client engagement. It is capability evidence and it is described as such.

The problem

Enterprise teams needed a reusable way to ask Arabic and English questions over private documents while keeping source control, predictable answers, and deployment flexibility. Those requirements normally conflict: the fastest route to good answers is a hosted model with a large context window, which is precisely what an organisation with private documents and infrastructure constraints cannot adopt. Each customer therefore rebuilt the same pipeline against its own language, cost, latency, and hosting limits.

What we built

We developed a document question-answering pipeline covering extraction, cleaning, chunking, embedding, vector search, prompt assembly, answer generation, and citations, with the model and retrieval components independently replaceable to suit a given customer's infrastructure.

What moved

Measure Before After
Information-retrieval time down 70%
Grounded-answer accuracy up 25%
Configuration time for a new document collection down 40%

† Measured against the prior level of the same measure. The source publishes the size of the movement, not the figure it moved from.

Information-retrieval time

Prior level 0 100 200 down 70%

Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.

Grounded-answer accuracy

Prior level 0 100 200 up 25%

Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.

Configuration time for a new document collection

Prior level 0 100 200 down 40%

Measured against the client's own prior process, indexed to 100. The source publishes the size of the movement, not the absolute figure it moved from.

Figures are drawn from the practice's own delivery records for the engagement named, measured against the process that preceded it. They have not been through third-party audit, and none is presented as an average across clients.

What it turned on

Treating preprocessing, metadata filtering, retrieval evaluation, and bounded prompts as independent controls - rather than leaning on a larger context window - is what allowed a weak component to be swapped for a particular language, document type, latency target, or deployment environment without rebuilding the pipeline around it.

Service line
Bilingual Document AI

Start here

Which of the four is yours?

Tell us the documents and the monthly volume and we will send the two closest records, with the proof behind each and the honest note where the match is partial.

You get a reply within one working day, from the engineer who would do the work - not a sales sequence.