The library since 1.0: a Slavic push, a Celtic axis, and canonical memory

18 July 2026 · Nabu news

Version 1.0.0 marked the library as whole; the six development phases since have made it wider. As of 2026-07-17 the catalog holds 172,189 documents / 4,308,814 passages across 38 registered, synced sources, with 633,137 dictionary entries on the reference shelf — up from 170,711 / 4,267,659 / 458,238 at the 1.0 census three days’ worth of syncing earlier. All figures below are read from the live catalog. This entry is not a release note: no version has been tagged since v1.0.0, and whether this accumulation warrants a v1.1 is a question that stays open.

The Slavic push. The reference shelf’s largest single arrival is the StarLing / Tower of Babel package (27,397 entries), ingested under a written grant from its maintainer with each database’s compilers credited on every surface: Pokorny’s complete Indogermanisches Etymologisches Wörterbuch (2,222 roots), Nikolayev’s Walde-Pokorny-based PIE database (3,291 etymologies with per-branch reflex columns), the Common Germanic and Baltic databases — and, for the Slavic axis, Vasmer’s etymological dictionary of Russian in the Trubachev edition, 18,239 entries, so define сигать now answers from the standard reference. Beside it, the Slovenian historical dictionary shelf (ZRC SAZU via CLARIN.SI, CC BY, 139,405 entries) unites Pleteršnik’s Slovene-German dictionary of 1894–95, the lexicon of Janez Svetokriški’s Baroque sermons, and the complete word inventory of sixteenth-century Slovenian print — closing the gold-lemma-to-dictionary loop for Slovenian. And the damaskini corpus (CC BY-SA) extends the Slavic shelves southward: 23 gold-annotated Balkan Slavic witnesses of the fifteenth through nineteenth centuries on the Church Slavonic–Bulgarian continuum, each with an aligned English translation, bringing Bulgarian into the gold-lemma index.

The Celtic axis. An entirely new wing, synchronized 2026-07-17. CorPH — the Corpus PalaeoHibernicum of the ERC ChronHib project at Maynooth — contributes 76 Early Irish documents of the seventh through tenth centuries (the Annals of Ulster, Vita Columbae, Blathmac, and the Milan, St Gall, and Würzburg gloss corpora) with 136,559 gold-lemmatized tokens: the library’s first Old Irish. The epigraphic record arrives from both sides of the Channel: the RIIG corpus brings 428 Gaulish inscriptions in Gallo-Greek and Gallo-Latin scripts with dated findspots and French translations, and Ogham in 3D some 500 Irish ogham stones in real Ogham codepoints with aligned transliteration layers (held under the restrictive non-commercial reading while the project’s two conflicting license statements await clarification). Old Irish, Middle Irish, and Middle Welsh Wiktionary extracts and two Old Irish treebanks from Universal Dependencies complete the axis. The gold-lemma index now answers in seventeen languages.

Canonical memory, at every grain. The metadata framework begun at 1.0 grew three limbs. Every registered source now carries a curated dossier — description, themes, key works — served on the new nabu list content census (one line per shelf; a full card per source) and checked against the public shelf map at every development gate by a mechanical drift rule. nabu note opens the notes shelf: the owner’s own annotations on any citable URN — scholia of one’s own — resolution-checked before anything is written, rendered wherever the target is shown, and served to AI clients under the target’s withholding rules. And nabu ingest now accepts URLs, downloading first and recording the given address in the manifest; the owner’s local-library shelf, empty by design at 1.0, has entered use.

Honesty notes. The ogham shelf’s license conflict is recorded verbatim in the source inventory and the shelf is held at nc pending a reply. The exact gold-lemma row count awaits a fresh census — the last verified figure, 2,852,069 rows, predates the Old Irish and Bulgarian layers. The numbers here will not stand still, and the next entry will carry the new ones.

← All news  ·  Atom feed