Capabilities
14 service lines, grouped by who buys them
What you actually get
Grouped by the function that buys it, not by the technology that builds it. Nobody has a computer-vision budget.
The four engagement stages answer how you buy: a diagnostic, then a roadmap, then a build, then oversight. The 14 lines below answer the other question - what actually gets built or improved - and they sit in five groups named after the function that signs for them: finance and procurement, operations, HR and shared services, customer and revenue, and compliance and governance. Each line carries what it touches, named explicitly, and the strongest delivered proof behind it, including a plain statement on four of them that the proof is partial or absent.
The partial markers are what make the strong lines believable. A catalogue where all 14 look equally strong reads as a catalogue where nothing was checked, and a sophisticated buyer reads it that way in under a minute.
How 18 services in the market became the 14 we will sell
less 2 we have no proof of any kind for: procurement and supply-chain AI, predictive maintenance
less 2 that are the audit and roadmap stages under another name: AI governance advisory, data readiness and platform work
Three lines with more behind them than the name suggests
Each of these is a category anyone can claim. What follows is the failure mode we hit and the number that moved.
AI Voice Agents
Five voice systems in production. Most firms selling conversational AI in this market are selling a chat widget.
Voice pays off wherever a high volume of repetitive enquiries can be answered from a system of record and the caller is often not at a desk - on a plant, in a vehicle, on a site, or out of hours. Delivered in energy and petrochemicals, telecommunications, banking and healthcare. The same shape applies directly to utilities, ports and logistics, facilities management and O&M, and government service desks - applies to, not delivered for, and the distinction is worth keeping visible.
The industrial anchor: a petrochemicals operator, where 64% of reception and switchboard contacts were handled without a person and time to a location, contact or facility answer fell from 2.9 minutes to 31 seconds. It took about three weeks to stand up, by reshaping an existing voice platform rather than building from scratch - which answers the how-long-before-this-is-real question before it is asked.
- Numbers spoken aloud across an authentication boundary are where voice agents actually die.
- A misheard digit does not degrade the answer, it ends the call - and the customer starts again with a human, which is worse than never having offered the agent. First-pass capture of spoken numbers went from 84% to 98.6%, and that single metric is the difference between a deployable agent and a demonstration.
Put your own numbers through it → - Gulf dialect is not one accent.
- Word error rate went from 31% to 12.7%, and - the harder number - the spread between dialects narrowed from 19 points to 5. A system that works well on average and badly on one region's speech has failed for a whole governorate, and an average hides exactly that.
Where this proof is partial. All five voice systems are off-sector for an industrial buyer. Use this as proof of engineering, not of industry relevance - running production Arabic voice AI for a national operator is a delivery claim, and it should be read as one.
AI Document Intelligence
Eight delivered systems, across invoices, purchase orders, contracts and settlements, delivery notes, submittals, scanned paper and bilingual documents. One capability applied to many document types, not a finance-department tool.
Coverage is the prize, not speed. On a telecom operator's partner settlements, verification against contract terms went from around 15% sampled to 100% - and the first full-coverage cycle surfaced 2.3 times the discrepancy rate the sampled process had been reporting.
State plainly what that multiple is and is not. It is not a claim about model accuracy. It is a measurement of what sampling was hiding, because discrepancies were never evenly spread across the partner book - and a sample assumes they are. The same shape repeats on accounts payable: sample to 100% three-way reconciliation, and validation time per invoice from 12 minutes to 45 seconds.
Extraction accuracy by input quality, and what happened to the remainder
- Accuracy split sharply by input quality, and we published the split.
- On scanned paper in a public health ministry, extraction ran at 96% on clean digital forms, 84% on poor scans and 61% on handwritten entries. Rather than averaging those into one flattering number, 14% of fields were routed to human review, each shown with the region of the page it came from. A vendor quoting a single accuracy figure for document extraction has either not measured it across real inputs or is choosing not to say.
- The system stopped re-catching the same error and started removing its source.
- Repeat mismatches from the same supplier fell 67% once correction notices went back automatically. That is a second-order effect and it is usually worth more than the extraction itself, because it reduces the volume rather than processing it faster.
Where this proof is partial. The proof is telecom and financial-services documents, not a construction subcontractor's invoice. The machinery is identical; we say which it was.
Semantic Search
Six delivered systems. The value here is not speed - it is surfacing an answer that would never have been found, and stopping one that would have been confidently wrong.
The engineering problem is not relevance ranking. It is that the words the asker uses are not the words the document uses: someone asks about retiring early while the statute says commutation of entitlement prior to the qualifying period. Closing that gap is the whole discipline.
On a social-insurance authority's statutory corpus - Arabic legal language, which is where general-purpose embeddings degrade most - correct provision in the top three went from 62% to 93% on exact statutory terms. The number that matters more is 34% to 89%, on queries phrased the way a member actually talks. That second figure is the vocabulary gap closing: questions that used to return nothing usable now return the right provision.
- A confident wrong answer is worse than no answer.
- A system that trusts embedding distance alone will return a provision that is topically similar and governs a different category of person entirely. Provisions that read as near-neighbours can carry materially different entitlements, so a relevance-refinement stage checks each candidate against the actual question: roughly 1 in 6 retrieved candidates were discarded before a member ever saw them. A companion system pushed downstream correction errors from 12% to 3% the same way - that last number is the cost of a confident wrong answer, measured directly.
Put your own numbers through it → - Permission models are the real work, and it runs on-premises where it must.
- An index that ignores repository-level differences will surface a document to someone entitled to see the folder but not the file. One of the six was built where nothing could leave the building at all, and that variant exists because a judicial body required it.
Where this proof is partial. The AI Workspace packages this capability plus the consolidation of the access points around it. Its own benchmark is a speed story, because there the answer is known to exist and the problem is finding it faster across too many places to look. That is a different problem from this line's, and we do not borrow its number to describe this one.
Figures are drawn from the practice's own delivery records for the engagement named, measured against the process that preceded it. They have not been through third-party audit, and none is presented as an average across clients.
The 14 lines
14 is the ceiling, and the list is only readable because of the five groups. If a fifteenth is proposed, something comes off.
AI for Finance & Procurement
AI Document Intelligence
Full write-up above4 delivered, 4 published hereclosest proof is from a neighbouring sector
Four thousand supplier invoices a month, and three people typing them into the ERP. Month-end waits for them.
Reads supplier invoices, purchase orders, contracts, delivery notes and material submittals. Extracts the fields, validates them against what the system already holds, and flags what disagrees rather than what merely looks unusual.
Proof. A Gulf telecom operator's finance function - sampled review at 15% replaced by 100% coverage; per-document review ~40 min → ~6 min
Where this proof is partial. The proof is telecom and financial-services documents, not a construction subcontractor invoice. The machinery is identical; say which it was.
Buys it: CFO, finance director, procurement director
Report & Reconciliation Automation
2 delivered, 2 published hereclosest proof is from a neighbouring sector
Somebody rebuilds that report by hand every month, and it is three weeks out of date the day it lands.
Replaces the recurring manual report - pulling from the systems of record, grouping and scoring what arrives continuously, and producing the thing a person currently assembles.
Proof. A Gulf fuel-retail operator - reporting lag ~3 weeks → under 24 hours, sample → full coverage
Where this proof is partial. Partial. No case yet where a specific recurring finance report was replaced end to end. The Gulf fuel-retail operator shape - continuous unstructured input becoming a grouped, scored dashboard - is the honest analogue, and should be offered as one.
Buys it: finance, operations, HSE, quality
AI for Operations
Semantic Search
Full write-up above6 delivered, 3 published here
Someone asks about retiring early. The system returns something that reads like an answer - and it governs a different category of member entirely. Nobody catches it until the wrong expectation is already set.
Finds the answer keyword search would never reach, because the asker doesn't know the document's own vocabulary for their situation - and, just as important, catches the answer that reads correctly but doesn't apply, before it reaches anyone. The value here is not speed. It is surfacing what would otherwise never be found, and stopping what would otherwise be confidently wrong.
Proof. A Gulf social-insurance authority - correct provision in top 3: 62% → 93%; queries phrased in the asker's own words rather than statutory terms: 34% → 89%; ~1 in 6 retrieved candidates discarded as confidently wrong before a member ever saw them. The value is not speed - it is surfacing what keyword search would never find, and catching what a naive match would confidently get wrong. A Gulf commerce ministry - correct activity in top 5: 96%; ranked first: 81%; downstream corrections at licensing: 12% → 3%
Buys it: operations director, HSE, engineering, legal, compliance
Workflow Copilots & Agents
3 delivered, 2 published hereclosest proof is from a neighbouring sector
They already have four systems open. Nobody wants a fifth place to go and ask.
An assistant inside the tools people already use that drafts, summarises, answers and completes routine steps - and, critically, one operating environment rather than another destination.
Proof. A Gulf tax authority - new use-case deployment time −40%, analyst throughput +35%
Where this proof is partial. The platform story is the strong one, and it is the argument for doing the foundations once. Lead with the structure, not the assistant.
Buys it: department heads, COO, IT
Document & Request Triage
3 delivered, 3 published hereclosest proof is from a neighbouring sector
It arrives in a shared inbox and goes to whoever has time.
Classifies what comes in and routes it - service requests, RFIs, claims, applications - with the basis for each decision recorded.
Proof. A Gulf telecom operator - compliance-phrase detection recall 94%; sample → 100% reviewed
Where this proof is partial. Partial, and say so in these words - the classification and routing machinery is in production; this specific application would be new. That sentence has never once cost a deal, and asserting otherwise would.
Buys it: IT, shared services, procurement
AI for HR & Shared Services
AI for Customer & Revenue
AI Voice Agents
Full write-up aboveThe caller has to read out an account number, and if the system mishears one digit the whole call is wasted.
Answers the phone in Arabic and English, completes the routine transaction rather than only answering questions, and hands over cleanly when it should.
Where this proof is partial. All five are off-sector for an industrial buyer. Use it as proof of engineering, not of industry relevance - "we run production Arabic voice AI for a national operator" is a delivery claim, and it should be said as one.
Buys it: customer service director, COO, shared services
AI Sales Agents
We push the same offer to everybody and measure whether they took it.
Works out what to offer which customer next, and why - recommendation and next-best-action with the reasoning visible rather than a score on its own.
Where this proof is partial. The proof is banking and retail. For an industrial buyer this line is relevant where there is a dealer network, a spares catalogue, or a renewals book - say so, and do not offer it where there is not.
Buys it: sales director, commercial, marketing
Service & Conversation AI
20 delivered, 6 published hereclosest proof is from a neighbouring sector
Routine enquiries consume the people who should be handling the difficult ones.
Text and web assistants that complete a service action rather than only answering, plus call summarisation and quality review on what still reaches a person.
Proof. A Gulf telecom operator - first-pass spoken-number capture 84% → 98.6%
Where this proof is partial. This is the most crowded line in the market - every competitor sells a chatbot. Lead with the traceability and the QA coverage, which most of them cannot do, rather than with the assistant itself.
Buys it: customer service director, COO, shared services
AI for Compliance & Governance
Bilingual Document AI
4 delivered, 2 published here
Your contracts are in Arabic, your invoices in English, and your HR letters are both on the same page. The tool you bought last year reads one of the three.
Document intelligence on mixed-script material handled natively rather than translated first - which is where global tooling degrades and where the practice's strongest technical evidence sits.
Proof. Our own build - classification consistency +30%, extraction coverage +28%, manual review time −60%
Buys it: every function; usually reaches us through legal, compliance or finance
Risk & Anomaly Detection
1 delivered, 1 published hereclosest proof is from a neighbouring sector
The one that does not belong looks exactly like the others until somebody adds it up.
Flags the transaction, claim, submission or reading that is out of pattern, with a reviewable reason rather than a score alone.
Proof. Our own build - our own build. No isolated production metric. Weakest line we list
Where this proof is partial. The weakest line we list. Do not build a pitch on it. It is a legitimate answer when a prospect raises it, and it belongs in a T1 scope rather than a proposal.
Buys it: risk, finance, audit, compliance
Cross-functional
Forecasting & Planning AI
5 delivered, 5 published hereclosest proof is from a neighbouring sector
The forecast is a flat assumption that hides a twenty-point swing.
Demand, consumption, spares and cash forecasting - built on history that has been cleaned first, because that is usually where the error actually lives.
Proof. A mid-market online fashion and homeware retailer, ~14,000 SKUs - forecast error ~42% → 19%, after repairing demand history
Where this proof is partial. The finding matters more than the number here. Both cases repaired data definitions before modelling, which is the practice's whole thesis in miniature.
Buys it: operations, supply chain, finance
Inspection & Safety Vision AI
8 delivered, 3 published here
There are ninety cameras and nobody is watching them.
Reads images and video for defects, PPE breaches and asset condition.
Proof. A beverage bottling operation, three high-speed lines - defect escape −84%, detection hours → under 2 min, returns −61%
Buys it: HSE, quality, operations
AI Training & Enablement
No delivered work of our own behind this line yet.
The licences were bought and nobody uses them.
Teaching teams to use what has been built - practical enablement and adoption programmes rather than generic AI awareness training.
Proof. None of our own.
It is the only line that needs no reference, only a curriculum. It is the lowest-risk thing the practice can sell cold, and it is a natural attachment to any delivered engagement.
Buys it: HR, L&D, department heads
Start here
Ask for the two lines closest to your situation, not all 14.
Describe the documents and the volume and we will send the two that fit, with the proof behind each and the honest note where there is one.
You get a reply within one working day, from the engineer who would do the work - not a sales sequence.