We build AI systems that sit on your own data and do a specific job: answer questions from a policy library, extract fields from invoices and delivery notes, draft a first-pass response a human then approves, or route a case to the right desk. Every engagement starts by picking one process and agreeing what a saved hour is worth, because an AI project without a baseline can never be shown to have worked.
Retrieval before generation
A model that answers from memory will invent things. A model that answers from your documents, and cites which paragraph it used, can be checked. We build retrieval first — chunking, embeddings, a hybrid keyword-and-vector index, and permission filters applied at query time — so an answer is always traceable to a source the reader is allowed to see.
- Answers cite the document and section they came from
- Permissions enforced at retrieval, so nobody sees a file they could not open themselves
- Arabic and English indexed together, so a question in one language finds evidence in the other
- A refusal when the corpus does not contain the answer, rather than a confident guess
Evaluation is the deliverable
Before anything ships we build a test set from your real cases — usually 100 to 300 questions with expected answers written by whoever does the job today. That set becomes the gate: every prompt change, model upgrade and index rebuild is scored against it, so you find out that a change made things worse in CI rather than in a complaint.
100–300
graded cases in the evaluation set before launch
Humans stay in the loop where it matters
Automation earns trust by degrees. We start with the system drafting and a person approving, measure the edit rate, and only widen the autonomy where the edit rate says it is safe. High-consequence steps — anything touching payment, contracts or personal data — keep a named approver permanently.
Governance that survives an audit
Prompt and response logging with retention you control, PII redaction before anything leaves your boundary, model routing so sensitive workloads can run in-region or self-hosted, and a written record of what data trained or grounded what. This is what the UAE Personal Data Protection Law and your own security review will ask for, and retrofitting it is far more expensive than building it in.