v1.4.0 + nabu-data v1.0.0 — the loop closes

1 August 2026 · Nabu news

Two releases, cut together, because they are one story.

nabu-data v1.0.0 is the library’s publishing arm made real: twelve datasets — Sanskrit form-to-lemma from DCS gold, the Tibetan Wylie fold and verb paradigms, a segmented Classical Tibetan slice with its boundary F1 published in-band, Greek metrical scansions anchored to Perseus CTS, the Han and kyūjitai orthography folds, the Aozora and Kanripo gaiji censuses, the Sabellic loanword table, the cuneiform value-to-sign table flattened from the Oracc Sign List, and — the release’s centerpiece — the first re-publication: stable anchors for ACTib’s segmented eKangyur, 461,301 rows tying every Derge Kangyur passage to its ACTib line by citation and content fingerprint, with the measured match census (99.51% letter-exact) in-band and the 2,245 divergent lines published side-by-side as a proofreading worksheet. Every dataset is plain CSV with a Frictionless Data Package manifest naming the exact upstream version, the producing code version, and a re-runnable recipe. CC BY 4.0, three share-alike carve-outs stated per dataset.

Nabu v1.4.0 is everything since v1.3.0 — six phases in six days:

The census at the tags: 974,972 documents / 68,381,456 passages in 125 language codes across 108 sources, 1.4 million dictionary entries, twenty-three research desks — and, for the first time, a data repository going the other way.

← All news  ·  Atom feed