v1.2.0 — the East Asian libraries and the Germanic wave
22 July 2026 · Nabu news
Release v1.2.0 spans three phases and the library’s largest growth between any two tags. At v1.1.0 (19 July 2026) the catalog held some 737,000 documents and 11.4 million passages; v1.2.0 closes at 801,175 documents / 28,176,484 passages across 82 sources — the corpus more than doubled in three days of phases.
Most of that is the East Asian step change. The Kanseki Repository
and the CBETA Buddhist canon made Literary Chinese the library’s largest
language (13.2 million passages), with traditional, simplified, and
variant spellings folded to one search skeleton, and the character desk
(nabu char) answering a Han character across four millennia —
structure, sound, and attestation on one card. Aozora Bunko brought the
Japanese public-domain library (3.0 million passages, ruby readings
preserved as annotations, kyūjitai reachable from modern forms).
The Germanic wave completed the desk across all three branches: the Old Norwegian treebanks and the Poetic Edda of Codex Regius, the Old Saxon Heliand parsed to the token, the Middle High German reference corpus in its diplomatic transcription, the Scandinavian runic corpus in five text lanes, and diachronic Icelandic through IcePaHC. The gold-lemma layer grew accordingly: 16.2 million gold rows in 28 languages as of the tag — Middle High German entered third and Icelandic fourth on the day they landed.
The instruments kept pace: a compact nabu status (detail one flag
away), the focus profile (scope the read surfaces to your own
research desks — a classicist need never scroll past the Japonic
shelves), whole-word search (--word), and stored-text snippets
everywhere — search results now always show the text as the manuscript
carries it, never the folded search skeleton.
Numbers above are read from the live catalog and dated; per-shelf detail is on The Library, licensing on Sources & Licensing. The release is citable via the versioned Zenodo DOI minted from this tag.