Skip to main content
Bilingual Knowledge-Graph Platform for 200 Years of Gulf Archival History
Case study
EducationWeb Application22 weeks

A Gulf-region research university digital-humanities centre

EducationWeb Application22 weeks

Thirty-three PDF volumes of 1622–1810 Gulf history — 1.94M words over 5,657 pages — existed only as unsearchable prose, alongside a parallel Arabic edition whose text layer was unrecoverable. No structure, no entities, no map, no way to cite a passage.

The challenge

What they came to us with.

The corpus was a set of thirty-three InDesign-exported PDF volumes covering Persian Gulf, Arabian and Persian history from 1622 to 1810 — 1,939,670 words across 5,657 pages. As documents they were readable. As research material they were inert: no records, no entities, no dates a machine could sort, and no way for a historian to cite a specific passage and have anyone else find it again.

A parallel Arabic edition of the same thirty-three volumes existed, and it was worse. The PDF text layer was glyph-scrambled and unrecoverable, so the Arabic scholarship was effectively invisible to every tool that could read the English.

The requirement was not a search box over PDFs. It was a structured, bilingual, citable knowledge graph — people, places, ships, commodities, events and the relations between them — with every claim traceable to the page it came from.

Our approach

How we built it.

Extraction is deterministic and font-aware rather than model-guessed: year headers are identified by weight and point size, body text by its own typographic signature, so segmentation is reproducible across all thirty-three volumes. The Arabic edition was recovered separately by page-level vision transcription at 200 dpi, yielding 8.36 million characters the original text layer could not surrender.

Entities are resolved through a seven-tier cascade — exact, normalised, transliteration-family, token-overlap, prefix, phonetic and edit-distance — backed by thirty-eight curated transliteration families and a hand-built gazetteer. Candidate merges are scored across name overlap, shared correspondents, shared places and year overlap, with explicit vetoes for conflicting residences and titles, and a learned Splink model behind it.

Events, relations and context facts are layered on top and written into Postgres 16 with PostGIS and pgvector. A FastAPI service exposes the graph; a bilingual Next.js 14 application renders it as an interactive Leaflet map, a correspondence network, a facsimile viewer and a cited research agent that is architecturally forbidden from paraphrasing a quotation — citations are sliced from character offsets or refused.

The outcome

What shipped. What changed.

Source text captured from 33 volumes

Before: 0After: 1.94M words
99.99% capture

Arabic edition records

Before: 524After: 3,985
7.6x, 1:1 with English

Entities after source-verified audit

Before: 3,469After: 2,687
0 dangling refs

Arabic search recall, bare-alef query

Before: 8After: 295
37x

Place mentions stuck at a meaningless 0.50 confidence

Before: 985After: 53
-95%

Over-span event clusters

Before: 66After: 0
Eliminated

Want the same outcome for your team?

Tell us where you are now. You'll get a fixed price in writing before any work starts.