Nabu is a piece of personal research infrastructure that gathers the world’s openly licensed digital corpora of antiquity — Homer and the Greek canon, the Latin classics, the documentary papyri of Egypt, the Latin inscriptions of the Roman empire, the Sanskrit tradition and the Pali canon, cuneiform tablets, three millennia of Egyptian sentences, the Masoretic Hebrew Bible with its Targums and Peshitta, the Hittite tablet corpus, the universal CDLI catalog of cuneiform, the Ugaritic tablets, a millennium of Syriac, the inscriptions of pre-Roman Italy and Sicily, the New Testament in up to fifteen parallel witnesses, the complete corpus of Old English poetry, the classical Chinese library with the Buddhist canon, the Islamicate library in Arabic and Persian, the Tibetan Buddhist canon, the medieval lyric of Galicia, the Japanese public-domain library, and the runestones of Scandinavia — into a single library on the scholar’s own disk. Everything is stored as plain files plus SQLite: it is searchable by word or by dictionary lemma, citable to the exact verse or tablet line, explicit about every text’s license, and rebuildable from its canonical sources at any time. Because the library also exposes a read-only Model Context Protocol server, the AI assistants a researcher already uses can search, quote, and cite the whole collection while remaining structurally unable to alter a letter of it.
The project takes its name from the Mesopotamian god of scribes, patron of the tablet house and divine custodian of Ashurbanipal’s library. It is neither a website nor a reading application: it is a pipeline and a database, operated from the command line, and designed to outlive the services it draws from.
Find your desk
The collection is organized as 25 research desks — scholarly hats, each gathering the shelves, instruments and terminal setup for one tradition. Find yours:
- The Classicist — Greek and Latin letters read whole, Homer to the late grammarians.
- The Romanist — the Latin-to-vernacular continuum, charter Latin to Roland to the troubadours.
- The Papyrologist-Epigraphist — reads what survives on stone, sherd, papyrus and tablet, lacunae and all.
- The Slavicist — Cyril and Methodius to the damaskini, canon to vernacular.
- The Germanicist — Gothic and Old English to the Norse sagas and the runestones, the word-hoard of all three branches.
- The Celticist — from Lepontic stones to the Old Irish glossators.
- The Italicist — the languages of pre-Roman Italy, Oscan to Etruscan to Raetic.
- The Comparative Indo-Europeanist — laryngeals, reflex chains, the long descent of words.
- The Biblical scholar — one text across Hebrew, Greek, Latin, Syriac, Coptic and English witnesses.
- The Hebraist — Masoretic vowels, Qumran consonants, the Aramaic of the Targums.
- The Syriacist — the Peshitta and the estrangela bookshelf.
- The Ethiopicist — Aksum to the scriptoria: Enoch, Jubilees, and the Geʿez Bible.
- The Arabist — the Islamicate library whole, Quran and hadith to falsafa and adab.
- The Hittitologist — Anatolia in cuneiform, KBo and KUB by tablet and line.
- The Assyriologist — Sumerian, Akkadian, Ugaritic, Hittite: the tablet world entire.
- The Egyptologist — hieroglyphs to Coptic, one language across four millennia of script.
- The Iranologist — the Avesta to the Achaemenid inscriptions, Old Iranian liturgy toward Middle Persian.
- The Indologist — Veda to sastra, the Sanskrit library and its instruments.
- The Buddhologist — the dharma across the Pali, Sanskrit, Chinese and Tibetan canons.
- The Tibetologist — the Land of Snows from Dunhuang documents to the complete Derge canon.
- The Koreanist — the hanmun state record and the first hangul vernacular.
- The Southeast Asianist — the Indic cosmopolis at its eastern edge: Sanskrit charters beside the vernaculars they seeded.
- The Sinologist — the classical Chinese written world and its phonological deep past.
- The Japanologist — Old Japanese song to the Sino-Japanese dictionary tradition.
- The Librarian — the owner’s own shelves: dossiers, library, notes, and the sources’ own records.
Each desk is a reader’s-eye view of the same collection; the full directory, with every desk’s member shelves and recipes, is Research axes.
The holdings, in brief
As of 19 September 2026, the catalog records 2,274,296 documents comprising 106,902,485 passages in 167 language codes — from proto-cuneiform tablets of the late fourth millennium BCE to Meiji-era Japanese. Classical Arabic is the largest language on the shelves (33.3 million passages), ahead of Literary Chinese (24.3 million). The newest additions are the granted sources — the complete secular Galician-Portuguese lyric under a written grant, and the Ras Shamra Tablet Inventory with the concordances Ugaritologists cite — beside the complete openMGH critical editions (153 volumes of medieval Latin), DÉRom’s Proto-Romance etymons, and the lect layer: historical-stage ladders, stage-aware search, and honest reconstruction display through the nabu-lects registry. Together with 1,998,589 dictionary entries across 114 reference shelves and 21.9 million gold-standard lemma annotations in 45 languages (a further 35.9 million machine-suggested annotations are carried at an honestly labelled silver tier). All figures on this site are read from the live catalog, never estimated, and carry the date on which they were read.
A survey of the collections is given on The Library page; the full attribution and licensing record is on Sources.
A single verse, many witnesses
One illustration, pasted from a live run (12 July 2026; trimmed lines are marked). A single Gospel citation rendered across the aligned witnesses — Greek, Latin, Gothic, Old Church Slavonic, Old English, and more — each carrying its license label. Since this run, the Sahidic and Bohairic Coptic New Testaments have joined the hub (13 July 2026), bringing the registered witnesses to fifteen:
$ bin/nabu align "MARK 2.3"
MARK 2.3 — New Testament (parallel witnesses)
13 of 13 witnesses attest this ref
greek-nt — The Greek New Testament [grc] license: nc
urn:nabu:proiel:greek-nt:6563
καὶ ἔρχονται φέροντες πρὸς αὐτὸν παραλυτικὸν αἰρόμενον ὑπὸ τεσσάρων.
latin-nt — Jerome's Vulgate [lat] license: nc
urn:nabu:proiel:latin-nt:10368
et venerunt ferentes ad eum paralyticum qui a quattuor portabatur
gothic-nt — The Gothic Bible [got] license: nc
urn:nabu:proiel:gothic-nt:37435
jah qemun at imma usliþan bairandans, hafanana fram fidworim.
marianus — Codex Marianus [chu] license: nc
urn:nabu:proiel:marianus:36421
Ꙇ придѫ къ немоу носѧште ослабленъ жилами. носимъ четꙑрьми.
wscp — West-Saxon Gospels [ang] license: nc
urn:nabu:proiel:wscp:102359
& hi comon anne laman to him berende, þone feower men bæron.
WEB (English) — Mark [eng] license: open
urn:nabu:eng-web:mrk:2.3
Four people came, carrying a paralytic to him.
… (Armenian, SBLGNT, Clementine Vulgate, and the four CCMH OCS witnesses trimmed)
Every result is a resolvable citation — a stable URN pointing into the local catalog — rather than an unverifiable quotation. The same discipline governs all of the tools: dictionary citations resolve to live passages, intertext hits name the shared phrase, and every surface carries its license class.
Principles
- License honesty. Every ingested text keeps its upstream license, recorded per document and displayed on every surface. Roughly 99% of documents are public-domain or attribution-class; non-commercial and no-derivatives materials are handled under correspondingly stricter rules. See Sources.
- Longevity over convenience. Upstream projects restructure, lose funding, and disappear. The canonical layer is plain files under git; all databases are derived and rebuildable; texts an upstream deletes are retained, honestly labelled. Storage is deliberately boring: files, git, SQLite.
- Citations, not summaries. The unit of work is the citable passage. Search results, dictionary entries, alignment rows, and machine-readable exports all carry stable URNs.
- Measured claims. Corpus numbers are snapshots of one live installation, dated where they appear.
Where to go next
- The Library — what is on the shelves: collections, periods, sizes, licenses.
- Tools — the command-line instruments, organized by scholarly task.
- Examples — worked walk-throughs for a classicist, a papyrologist, a slavist, a comparativist, an assyriologist, a hittitologist, and a biblical scholar.
- About — what this project is, who it is for, and how it is built.