Six shelves in a day: the Rabbinic library, Geʿez, and Italy's stones

26 July 2026 · Nabu news

Some phases follow a single thread; this one pulled six at once, and all of them held. Planned in the morning from three research scouts, dispatched at ratification, merged by evening, synced and wired the next day — the library grew from 813 thousand to 947,273 documents / 66.8 million passages in the process.

The Rabbinic wave is the anchor. The Sefaria shelf — until now the Targums alone — opens onto the Rabbinic core: the Mishnah in three complete Hebrew editions (the Kaufmann manuscript among them), the Wikisource Babylonian Talmud in Aramaic, Guggenheimer’s Jerusalem Talmud, the Tosefta after codex Vienna, the Minor Tractates, and the Davidson/Steinsaltz translation lane carried at its honest nc tier. Talmud citations mint in the reference system scholars actually use — tamid:he:wikisource-talmud-bavli:25b.1 is daf 25, amud b, line 1, and Tamid really does start at 25b. Mishnah and Tosefta enter as Tannaitic Hebrew (hbo), the gemara as Aramaic (arc), both byte-faithful under the library’s normalization exemption.

The Geʿez shelf is the phase’s deep cut: the first Ethiopic holdings, from the Beta maṣāḥǝft transcriptions in Hamburg — the Ethiopic Bible at verse grain beside 1 Enoch and Jubilees, books that survive complete only in Geʿez — with Dillmann’s 1865 Lexicon linguae aethiopicae (13,727 entries) and the TraCES corpus’s 75,440 morphologically analyzed tokens, each one lemma-linked into Dillmann’s own entries. A twenty-second research desk, the Ethiopicist’s, was minted for it.

Around those two: EDR brings the inscriptions of Italy — 115,590 records, the geographic complement of Heidelberg’s provinces — in one checksummed archive; the Oracc CC0 pack grows the cuneiform shelf by 64 projects, Urartian entering gold-lemmatized through eCUT beside the Amarna letters and the Assyrian provincial archives; ReN adds Middle Low German — 1.49 million gold-annotated tokens of Hanseatic prose, straight into the gold index’s top five; and the comparativist desk completes its loanword layer with WOLD (64,289 lexemes with donor languages), the CLICS³ colexification network, and the StarLing Kartvelian base under the project’s standing grant.

The gold-lemma index now answers in thirty-seven languages — 19.2 million rows; the reference shelf passed one hundred dictionaries and 1.39 million entries. And because every card deserves a curator: all 305 held language codes now carry a written dossier — including the honest ones, like the code that turned out to be Kali’na of the Guianas rather than Anatolian Carian, and says so.

Local, license-honest, rebuildable — now from Sumer to the scriptoria of Aksum.

← All news  ·  Atom feed