The library since 1.0: a Slavic push, a Celtic axis, and canonical memory
18 July 2026 · Nabu news
Version 1.0.0 marked the library as whole; the six development phases since have made it wider. As of 2026-07-17 the catalog holds 172,189 documents / 4,308,814 passages across 38 registered, synced sources, with 633,137 dictionary entries on the reference shelf — up from 170,711 / 4,267,659 / 458,238 at the 1.0 census three days’ worth of syncing earlier. All figures below are read from the live catalog. This entry is not a release note: no version has been tagged since v1.0.0, and whether this accumulation warrants a v1.1 is a question that stays open.
The Slavic push. The reference shelf’s largest single arrival is the
StarLing / Tower of Babel package (27,397 entries), ingested under a
written grant from its maintainer with each database’s compilers credited
on every surface: Pokorny’s complete Indogermanisches Etymologisches
Wörterbuch (2,222 roots), Nikolayev’s Walde-Pokorny-based PIE database
(3,291 etymologies with per-branch reflex columns), the Common Germanic
and Baltic databases — and, for the Slavic axis, Vasmer’s etymological
dictionary of Russian in the Trubachev edition, 18,239 entries, so
define сигать now answers from the standard reference. Beside it, the
Slovenian historical dictionary shelf (ZRC SAZU via CLARIN.SI, CC BY,
139,405 entries) unites Pleteršnik’s Slovene-German dictionary of
1894–95, the lexicon of Janez Svetokriški’s Baroque sermons, and the
complete word inventory of sixteenth-century Slovenian print — closing
the gold-lemma-to-dictionary loop for Slovenian. And the damaskini
corpus (CC BY-SA) extends the Slavic shelves southward: 23
gold-annotated Balkan Slavic witnesses of the fifteenth through
nineteenth centuries on the Church Slavonic–Bulgarian continuum, each
with an aligned English translation, bringing Bulgarian into the
gold-lemma index.
The Celtic axis. An entirely new wing, synchronized 2026-07-17. CorPH — the Corpus PalaeoHibernicum of the ERC ChronHib project at Maynooth — contributes 76 Early Irish documents of the seventh through tenth centuries (the Annals of Ulster, Vita Columbae, Blathmac, and the Milan, St Gall, and Würzburg gloss corpora) with 136,559 gold-lemmatized tokens: the library’s first Old Irish. The epigraphic record arrives from both sides of the Channel: the RIIG corpus brings 428 Gaulish inscriptions in Gallo-Greek and Gallo-Latin scripts with dated findspots and French translations, and Ogham in 3D some 500 Irish ogham stones in real Ogham codepoints with aligned transliteration layers (held under the restrictive non-commercial reading while the project’s two conflicting license statements await clarification). Old Irish, Middle Irish, and Middle Welsh Wiktionary extracts and two Old Irish treebanks from Universal Dependencies complete the axis. The gold-lemma index now answers in seventeen languages.
Canonical memory, at every grain. The metadata framework begun at 1.0
grew three limbs. Every registered source now carries a curated
dossier — description, themes, key works — served on the new
nabu list content census (one line per shelf; a full card per
source) and checked against the public shelf map at every development
gate by a mechanical drift rule. nabu note opens the notes shelf:
the owner’s own annotations on any citable URN — scholia of one’s own —
resolution-checked before anything is written, rendered wherever the
target is shown, and served to AI clients under the target’s withholding
rules. And nabu ingest now accepts URLs, downloading first and
recording the given address in the manifest; the owner’s local-library
shelf, empty by design at 1.0, has entered use.
Honesty notes. The ogham shelf’s license conflict is recorded
verbatim in the source inventory and the shelf is held at nc pending a
reply. The exact gold-lemma row count awaits a fresh census — the last
verified figure, 2,852,069 rows, predates the Old Irish and Bulgarian
layers. The numbers here will not stand still, and the next entry will
carry the new ones.