
A publishing-adjacent services company handling 500+ DOCX assets per month
Editorial ops was running 11 separate tools to move a document from intake to edit to merge to format to export. Turnaround was unpredictable and any one tool outage stalled the whole pipeline.
What they came to us with.
Editorial ops was orchestrating 11 separate tools — intake in one system, editing in another, merges in a third, PDF export in a fourth, and a spreadsheet that tied them together. Turnaround was unpredictable and any one tool outage stalled the whole pipeline.
Big volume merges were the worst pain point: synchronous operations against multi-hundred-page DOCX files would OOM and time out unpredictably, and failures were not cleanly recoverable. Editors were keeping manual spreadsheets to track what had and hadn’t gone through.
How we built it.
We replaced the 11-tool chain with a single Next.js 16 workspace. Editors drag and reorder document sections with dnd-kit, convert between DOCX, PDF, and HTML (Aspose Words Cloud + docxtemplater + mammoth), collaborate inside resizable side-by-side panels, and operate inside role-gated review flows. Authentication, storage, and access control are Firebase; errors flow to Sentry.
The architectural unlock was moving every document merge onto Google Cloud Tasks with idempotent task IDs. Synchronous merges could not survive large volumes; async made failures recoverable, kept the UX feeling instant, and let us scale heavy operations independently of the request path.
What shipped. What changed.
Tools in editorial pipeline
Median doc turnaround
P99 merge failure rate
Review cycle method
Keep reading.

Education — A Gulf-region research university digital-humanities centre
Bilingual Knowledge-Graph Platform for 200 Years of Gulf Archival History
Thirty-three PDF volumes of 1622–1810 Gulf history — 1.94M words over 5,657 pages — existed only as unsearchable prose, alongside a parallel Arabic edition whose text layer was unrecoverable. No structure, no entities, no map, no way to cite a passage.
- 99.99% source-text capture across 33 volumes — 1,939,480 of 1,939,670 words, every drop classified as page furniture
- A knowledge graph of 2,687 entities, 16,377 events and 19,814 typed relations drawn from 4,328 records spanning 1622–1810
- Arabic edition taken from 524 to 3,985 records at 1:1 parity with English, with only 1.9% machine-translated

Education — An Arabic-language education initiative
Research-Grade Linguistic Analysis Platform
A corpus of 676 Arabic doubled-verb conjugations needed to be explorable by researchers and students alike. Existing academic tools required CSV wrangling and produced static PNGs that nobody could interact with.
- All 676 entries browsable in a single interactive view
- Phonetic features auto-discovered from sifat strings (no hand-coded feature set)
- Silhouette-guided k-selection removes the need to hand-tune cluster count
Want the same outcome for your team?
Tell us where you are now. You'll get a fixed price in writing before any work starts.