
A MENA-region AI startup shipping two consumer-facing products
Two AI products (one retrieval-heavy, one vision-heavy) were being built as separate backends. Duplicated auth, duplicated RBAC, duplicated observability, and diverging fast.
What they came to us with.
The startup was standing up two products — one retrieval-heavy, one vision-heavy — as separate FastAPI backends. Auth was duplicated, RBAC was duplicated, observability was duplicated, and the two codebases were already diverging in subtle ways. Any cross-cutting fix needed to be applied twice.
The instinct on day one was microservices. The reality was a four-person engineering team that could not afford the operational overhead of two deploys, two auth systems, and two observability stories.
How we built it.
We consolidated onto a modular monolith: one FastAPI app, 11 route groups, 36 endpoints, one set of migrations. RS256 JWT auth (10-minute access, 30-day refresh) with token versioning so a password change invalidates every session without growing a revocation list. 4-role RBAC enforced via FastAPI dependency injection. Prometheus metrics and OpenTelemetry spans are emitted across every route.
The key containment pattern: per-product routers and per-product ARQ workers. Shared infra, isolated blast radius. A slow vision job cannot starve the retrieval product, and a bad migration in one product’s models cannot take the other down. Deployed as Docker multi-stage images to AWS ECR (me-central-1), fronted by Traefik.
What shipped. What changed.
Shared infra components
Time to add new product surface
Observability coverage
Blast-radius isolation
Keep reading.

Finance — A Gulf-region tax and compliance advisory firm
RAG Assistant for MENA Regulatory Compliance
Advisory staff were answering the same 40-50 recurring UAE corporate-tax questions by hand, each pulling 2-3 regulatory PDFs. Turnaround averaged 6 hours and junior staff frequently missed cross-references between VAT and CT documents.
- Avg. first-answer latency under 4s end-to-end
- Retrieval precision@5 improved from 0.61 to 0.88
- L1 staff resolution rate climbed from 35% to 82%

Retail — A UAE home-furnishing retailer with a 40k-SKU catalogue
Visual Product-Discovery Assistant for Home-Furnishing Retail
Customers shared Pinterest-style inspiration photos over WhatsApp and expected matching SKUs in return. Manual matching cost 20-30 minutes per enquiry; most customers dropped off before the retailer could respond.
- Enquiry-to-shortlist time compressed from ~25 minutes to ~45 seconds
- 58% of enquiries now resolve without staff involvement
- Rate-limited (5 req / 5s burst, 20 req / 60s sustained) to cap SerpAPI spend
Want the same outcome for your team?
Tell us where you are now. You'll get a fixed price in writing before any work starts.