Arabic — The Arabist
The Arabist — the Islamicate library whole, Quran and hadith to falsafa and adab.
The OpenITI lane: premodern Arabic and Persian literature at corpus scale — Quran and hadith, history and biography, law and falsafa, the dīwāns and adab — with the Persian shelf (Ḥāfiẓ, Ibn Sīnā) riding the same Arabic-script fold that makes ara/fas cross-searchable (P41-3) — and DiCCAS, disaster accounts excerpted from ten classical sources with catastrophe terminology tagged — and, since P96, the KITAB text-reuse instrument: upstream-computed reuse edges joining the held OpenITI books pairwise on the intertext desk.
The premodern Islamicate library is LIVE: OpenITI’s ~9,079 primary
texts (~1.12 B words, the 2026 release zip md5-pinned into canonical)
made Classical Arabic the largest language on the shelves — 33.3
million passages — with the Persian lane (fas) riding the same
corpus. Holdings below are read live.
New here? The Quickstart sets up the library in minutes.
The shelves
A source wears every desk it serves — these four answer this desk. Holdings are read live from the catalog and dated; a shelf with nothing synced yet says so.
| Source | Holds | License | Status | Holdings (as of 19 September 2026) |
|---|---|---|---|---|
kitab |
feature module | nc | wired · manual | nothing held yet |
openiti |
texts | nc | wired · manual | 9,079 docs / 34,631,499 passages |
diccas |
texts | nc | wired · manual | 10 docs / 879 passages |
kitab-reuse |
feature module | nc | wired · manual | nothing held yet |
Languages on this desk (live doc-or-entry counts as of 19 September 2026): ara 8,738 · fas 351.
The desk’s instruments
- No gold lemmas — OpenITI is unannotated. The corpus carries no
morphology or lemma layer, so
--lemma,vocabandformulasdo not apply to this desk; its instruments are full-text search across the whole Islamicate shelf and the timeline. - The ara/fas Arabic-script fold (P41-3): one search skeleton across
the ی/ي and ک/ك keyboard split, maqsura, taa marbuta, tashkeel, tatweel
and ZWNJ — so a query typed on either keyboard reaches the stored form
whichever keyboard wrote it. Search-side only; the stored bytes stay
pristine (conventions §9), and
--lang ara/--lang fasscope to one shelf. - The AH death-year timeline: every OpenITI urn opens with the
author’s 4-digit hijrī death year, so the OpenitiDates extractor lands
each text on the calendar as a CE terminus — round(AH × 0.970225 +
621.57), the tabular conversion — and
--from/--toand--centuryscope the shelf by when its authors died (Ḥāfiẓ d. AH 792 = 1390 CE). No gazetteer, so no--placeon this desk. - License posture —
nc(CC BY-NC-SA 4.0, the Zenodo record’s only grant): the shelf is MCP-excluded, so the AI server never serves OpenITI passages; the CLI reads them for local research.
Working the arabic desk
The generic axis surfaces — every desk answers to these, in working order (enable once, sync, then query):
nabu enable arabic # first time: put this desk's shelves in this box's profile
nabu sync arabic # fetch/refresh the desk's enabled members
nabu list --axis arabic # the shelf census, this desk only
nabu axis arabic # the desk card: members, holdings, gold coverage
nabu search WORD --axis arabic # a query scoped to this desk's shelves
This desk’s own surfaces:
nabu show urn:nabu:openiti:0792Hafiz.Muntasab.PDL00074-per1 # Ḥāfiẓ's Muntasab — Persian verse in the %~% hemistich notation
nabu search الله --lang ara # Allāh across the Arabic hadith, dīwān and falsafa shelves
nabu search دانی --lang fas # the cross-keyboard ی/ي fold (P41-3) — an Arabic-yeh query (U+064A) still finds Ḥāfiẓ's farsi-yeh دانی (U+06CC)
nabu search --from 1300 --to 1400 --axis arabic # term-less browse of the 14th-c.-CE shelf on the AH death-year timeline (live-verified 23 July 2026)
Ask your model
With the MCP server connected, this desk answers conversational research questions. Each example ran live against this library:
- “Who reused this passage of Ibn Sīnā’s Šifāʾ?” →
nabu_links urn:nabu:openiti:0428IbnSina.ShifaAfcalWaInficalat.ALCorpus00001-ara2:1.223.4— kind=”reuse” edges from the KITAB pairwise data: the passage reappears in Faḫr al-Dīn al-Rāzī’s Mulaḫḫaṣ (two centuries later) with character/word offsets riding the edge detail — upstream- computed alignments, distinct from nabu’s own intertext detection.
Terminal setup
- Arabic and Persian (ara/fas) are RTL and reuse the hbo/arc
machinery:
isolates: truewrapping, the same modes, the same honesty footer. The terminal owns the direction — the iTerm2 ≥ 3.6.0 RTL toggle (Settings → General → Experimental; Terminal.app has no bidi at all). - Shaping stays degraded even with bidi on. Arabic is a connected
script and a cell-grid terminal cannot fully join it, so what you get
is right-to-left, legible, unligatured Arabic — fine for scanning
search hits and citations, not for sustained reading (use
nabu exportand a real text view for that). Font:font-noto-naskh-arabic— naskh stays legible at terminal sizes; a dedicated iTerm2 reading profile with it boosted a few points is the workable setup (docs/display.md §2). - Nothing to strip, and no
--display translit. The P41-g census found consonantal standard-block text only (no tashkeel), sodefault,plainandfullrender the same bytes (isolates aside); and Arabic romanization is deliberately not built — the standards conflict and unpointed text lacks the vowels a romanization needs, so ara/fas pass through the translit mode unchanged (docs/display.md §1d).
The full guidance, per script, is on the display page.
One of the twenty-five research desks; the flat shelf map is The Library and the reasoning is docs/axes.md.