News
One entry per release or development phase: what entered the library, what the tools can now do, and the honest numbers as of each date. Entries are added at every phase gate (the same pass that re-syncs the rest of this site from the repository documentation). Subscribe by Atom feed.
- 18 September 2026 v1.6.0 β four axes, and a public data wave β The release that completes the library's fourth document axis and sends the layers out into the world: classification across two million documents, the unified Layers page, a clean health board, and the largest nabu-data publication wave since the repository opened.
- 18 September 2026 Layers, and the literary family β The classification tree learns depth β literature becomes one family with poetry, narrative, drama and their siblings inside it β the unmapped worklist shrinks to a quarter of its size, and the site's three axis pages merge into one Layers page.
- 14 September 2026 The classification axis: what kind of document is this? β The library gains its fourth axis. Beside when, where, and in what language variety, every document can now answer what KIND of text it is β funerary, administrative, letter, divination, historiography β one ruled vocabulary folded over every upstream jargon, with the original label preserved verbatim beside the fold.
- 11 September 2026 The instruments phase: the library learns places and people β Three research instruments enter the library β a historical-China gazetteer, a 658,000-person biographical database, and the KITAB text-reuse graph β and the first text-mining lane runs: 3.77 million place-name candidates over the Chinese canon, plus a dictionary card that now says when a word was first written down.
- 3 September 2026 The search phase: six dragons, a 51-gigabyte diet, and an attic β One phase, all of it about finding things: exact Han-character search over the 23-million-pair CJK index, a full-text engine rebuilt to half its size with answers unchanged, a search view for everything the library ever withdrew β and the semantic-search machinery built end to end, awaiting its first overnight build.
- 2 September 2026 The Southeast Asia desk β and the first Old Mon text anywhere β The library's 25th desk opens whole in one phase: the DHARMA epigraphic corpora (Old Khmer, CampΔ, Nusantara, Pyu), the kakawin library, the Old Burmese inscriptions of Bagan, and the Old Javanese Wordnet β ten languages never held before, among them two inscriptions of Old Mon, a language with no digital corpus anywhere until now.
- 1 September 2026 A million records of Korea β the library doubles in a day β The historical slice of the Open Korean Historical Corpus lands: 1,198,779 records β the Samguk sagi of 1145, the Ilseongnok court diaries, the munjip mass of the Korean literati β each carrying its own license label. The library's document count more than doubles in a single source.
- 31 August 2026 The Iguvine Tables β and a hundred million passages β The seven bronze tables of Iguvium β the longest ritual text to survive from pre-Roman Italy, and the corpus of the Umbrian language β arrive with the Oscan inscriptions in the TITUS edition, under an extended personal grant. The same week's census crosses one hundred million passages, and the rebuild machinery learns to trust its own stamps.
- 31 August 2026 The poets of early Akkadian β SEAL joins by grant β Sources of Early Akkadian Literature β the edition of record for Old Babylonian Gilgamesh and the oldest Akkadian poetry β enters the library under its editors' written personal-research grant: 408 compositions at line grain, damage brackets and all. The Coptic dictionary gains a second etymological witness the same week.
- 30 August 2026 The rhyme books arrive β an Old Mandarin desk, and the eastern shelves widen β Seven centuries of Chinese phonology land as a connected series of rhyme books β the Guangyun and Wang Renxu's Qieyun, Fujita's restoration of the Qieyun itself, the 'Phags-pa θε€ει» and the δΈει³ι» of 1324 β beside a 1.9-million-passage ClassicalβModern parallel corpus, the Monlam Tibetan lexicon, and the first annotated Newar shelf.
- 27 August 2026 One million documents β the presses of England open the door β EEBO-TCP Phase I lands: 25,368 hand-keyed early-modern English texts β Milton and Raleigh beside the sermons, broadsides and civil-war tracts that carry the period β and the library crosses one million documents, closing in on ninety million passages.
- 26 August 2026 The granted doors β the sagas and the Gawain-poet arrive β Two doors opened by correspondence now stand as shelves: the Menota archive's 91 medieval Nordic manuscripts β LaxdΓ¦la saga, the Codex Wormianus, the Old Norwegian homily book β and the Corpus of Middle English, the quotation base behind the Middle English Dictionary. The library crosses 75 million passages, and the registry learns Norwegian and Danish.
- 23 August 2026 v1.5.0 β and the registry gets a version of its own β The release that closes the arc: nearly a million documents, 73.6 million passages, and time as a first-class axis β every source's own date forms turned into machine bounds, a chronological browse, and 128,051 documents newly placed on their language's timeline. Beside it, the lect registry cuts its own first release: nabu-lects v1.0.0, with dialects given a proper home.
- 21 August 2026 The annals, the fathers, and the honest instruments β Three phases in one arc: the Korean desk opens with half a millennium of Joseon state records, the instruments get honest β search stops inventing Greek fragments, erased papyrus text returns in double brackets β and the Patrologia Latina brings the church fathers home beside Catalan print, Portuguese drama, and the small Romance voices.
- 18 August 2026 Doubt made visible, and the vernaculars arrive β Two waves: the uncertainty doctrine starts rendering β certain is silent, deviation is labeled, the upstream word rides verbatim β and the medieval vernaculars enter the library: Old Spanish with its lenguas resolved, the Commedia, Old Swedish law, Achaemenid Babylonian, and a derivability proof that finally ran green.
- 12 August 2026 The layers learn to join β The Composition Phase: search becomes identity-aware about places, gains writing-system and geo-radius cuts, searches by cuneiform sign, and grows a deterministic filter-first mode β the held layers finally answering one question together.
- 11 August 2026 The library proves it can die β then publishes its layers β Three phases in one arc: the derivability law makes everything permanent live in three plain-file folders and proves it by an actual restore drill; the character desks grow a graded-reading lane and a didactic overlay from the scribal school next door; and nabu-data ships eight new datasets β the library's stage, dating, place, and sign layers, published whole, with 95,054 cuneiform tablets upgraded to single-year dating on the way out.
- 9 August 2026 One glyph in, everything out β sign cards for Cuneiform and Egyptian β nabu char grows beyond the Sinoverse: a cuneiform sign card re-joined entirely from data already on the shelf (the Oracc Sign List and the CDLI glosses riding inside it), and an Egyptian card anchored on Unicode 17's Unikemet file β 5,067 hieroglyphs with descriptions, functions, values, and three catalog concordances. Both cards count their signs in the wild, across millions of held passages.
- 9 August 2026 Places become decisions β the third dimension lights up β Two phases of the places program land: the gazetteers of the ancient world held locally under one namespaced index, a sister registry that records which place each source's name-string actually means, and a 3.9Γ jump in place-linked documents β 151 thousand to 585 thousand. Plus a new pattern for data only a human can download, and the registry's license to mint places of its own.
- 7 August 2026 The dating pendings close β three gaps, three different causes β Wave two of the core-layers plan: the three biggest dating gaps in the library close, and each turns out to have a different real cause than the survey guessed β a directory one level deeper than the walker looked, sibling documents that never inherited their stone's date, and a period vocabulary nobody had banded. Plus one ruled table now answers "when is Ur III?" for every consumer at once.
- 7 August 2026 Every source answers for its layers β The core-layers frame lands: one posture file where all 77 sources answer for dating, places and script; the script projection goes live (search a corpus by the script its texts are actually written in); the artifact-script field marks where the original differs from the held page; and two new standing audits β one of which caught twelve dangling gazetteer refs on its first run.
- 6 August 2026 The script axis β what script is this text actually in? β The lect identifier gains a fifth axis: ~script claims the writing system of the text as held, machine-checkable against the bytes β and a measurement of the catalog's own script-suffixed codes is what overturned the registry's earlier doctrine. Plus: four Sefaria documents recovered from quarantine by teaching the parser nested schema nodes, and language dossiers gain their stage ladders.
- 6 August 2026 The lect layer stabilizes β and becomes part of the core β The phase after load-bearing is housekeeping with teeth: the reversed-interval defect cluster that fed the date inference bogus bounds is repaired across six sources, the deferred stage mappings close (New Kingdom Egyptian, Qumran Aramaic, the Sefaria Talmud/Targum split, Standard Babylonian as a composed register), every query surface speaks --lect, and every source in the library now carries a recorded lect posture β identity being an honest answer.
- 4 August 2026 The lect layer goes load-bearing β 408,884 stage assignments β One phase after the nabu-lects registry arrived, the library stops merely reading it and starts ruling with it: a rebuild-proof journal of per-document stage assignments, compiled from the catalog's own period labels and per-inscription dates β Sumerian, Akkadian, Latin epigraphy, Egyptian and Vedic Sanskrit staged at scale, every assignment carrying its evidence and audited against the documents' own dates.
- 1 August 2026 v1.4.0 + nabu-data v1.0.0 β the loop closes β A synchronized double release: Nabu v1.4.0 β the signs desk, the Tibetan consumers, the granted sources β and the first tagged release of nabu-data, twelve published datasets with full derivation provenance, three of them consumed back by the library that produced them.
- 28 July 2026 The fourth canon: Tibet takes the twenty-third desk β Phase 48 lands the complete Derge Kangyur and Tengyur, 84000's English layer folio-paired with the Tibetan, the Old Tibetan documents of Dunhuang, and the MahΔvyutpatti bridge β and the Buddhist desk holds all four canons at once.
- 28 July 2026 Elephantine: one island, four thousand years β Phase 47 crawls the Berlin Elephantine archive β the Judean garrison's Imperial Aramaic beside Greek, Demotic, Hieratic and Coptic β and turns the incident trail into infrastructure: every catalog lane now refreshes at sync and is guarded by health checks.
- 26 July 2026 v1.3.0 β from Sumer to the scriptoria β The v1.3.0 release: four phases since v1.2.0 β the Islamicate library, the meter layer and the place index, the Romance continuum, and the Rabbinic and GeΚΏez shelves β 66.8 million passages across twenty-two research desks, gold lemmas in thirty-seven languages.
- 26 July 2026 Six shelves in a day: the Rabbinic library, GeΚΏez, and Italy's stones β Phase 46 opens the Rabbinic corpus at daf-grain citation, brings the first Ethiopic shelf (Enoch and Jubilees, complete only in GeΚΏez), the inscriptions of Italy, Urartian, Middle Low German, and the comparativist loanword layer β six packets planned, built, synced and wired inside twenty-four hours.
- 25 July 2026 The Romance phase: Latin's daughters take a desk β Phase 45 builds the Latin-to-vernacular continuum as a research desk of its own: the MGH critical editions, late-antique prose, Old French from the Serments de Strasbourg to the fifteenth century, and the Romance treebanks β while the meter layer doubles and place lookups drop from seconds to milliseconds.
- 24 July 2026 The instruments phase: meter, places, and asking your model β Phase 44 builds instruments over the library: a metrical scansion layer for Latin and Greek verse, a place desk joined to the Pleiades gazetteer, twenty million tokens of silver-annotated Greek, Croatian Latin β and, for AI assistants, an eleventh MCP tool plus live-verified example questions on every desk page.
- 23 July 2026 The efficiency phase: the library rebuilt, and every command flies β Phase 42 rebuilds the doubled library from canonical sources in under five hours and retires every slow command: status drops from four minutes to two seconds, vocabulary profiles from eighteen seconds to half a second, and searches for the commonest words in the corpus answer instantly with an honest footer.
- 22 July 2026 The Arabic phase: the Islamicate library, staged β Phase 41 opens the Arabist's desk on OpenITI β premodern Arabic and Persian at corpus scale, the largest single corpus the library has ever staged (~9,106 texts / ~1.12 B words), read through a bespoke mARkdown parser, an AH death-year timeline, and a shared Arabic-script search fold that makes Arabic and Persian cross-searchable.
- 22 July 2026 v1.2.0 β the East Asian libraries and the Germanic wave β The corpus more than doubles between tags: the classical Chinese library and the CBETA canon, the Japanese public-domain library, the character desk, and all three Germanic branches β with the compact status display, the focus profile, and whole-word search.
- 22 July 2026 The Germanic phase: from two languages to all three branches β Phase 40 widens the Germanicist's desk from Gothic and Old English to all three Germanic branches β Old Icelandic, Old Norwegian and the Poetic Edda, the Old Saxon Heliand, Middle High German, and the runic inscriptions β through two new parser families, a sibling, a custom reader, and one registry line.
- 22 July 2026 The registry phase: what a thing is, and what it costs β Phase 39 teaches the library to say what each registered thing is β source, shelf, or module β makes rebuild stamps language-aware, finds two real performance bugs by their names, and recovers all 1,191 quarantined Aozora works in one census.
- 21 July 2026 The Japanese phase: Aozora Bunko and the honest-glyph ladder β Phase 38 opens the Japanese reading desk β an adapter for Aozora Bunko's 17,000-work public-domain library β and replaces the blank rare-character placeholder with a four-rung display ladder that shows the most faithful rendering the evidence supports.
- 21 July 2026 The character desk: the largest shelf gets its instruments β Phase 37 equips the Chinese collection: a character card spanning four millennia per glyph, structure search by radical and component, the traditional-simplified fold, honest gaiji rendering, and the Daozang joining the Kanseki Repository waves.
- 20 July 2026 The engine phase: rebuilds measured, stamped, and made incremental β Phase 36 gives the derived layer its instruments: a stage profiler in every rebuild, derivation stamps that let unchanged sources skip re-derivation, and the query constants recalibrated to the settled 24.4-million-passage corpus.
- 20 July 2026 The atlas phase: research axes, personas, and honest surfaces β Phase 35 stops widening and maps: eighteen research axes with their personas over the eighty-source registry, axis-aware listing and syncing, and a standing audit that keeps the code's assumptions honest as the corpus grows.
- 20 July 2026 The Chinese libraries: Kanripo, the TaishΕ canon, and a concept lexicon β Phase 33 gives the Sino axis its bookshelf: the Kanseki Repository's classics, histories, masters and literature; CBETA's TaishΕ and Xuzangjing at print-citation grain; and the Thesaurus Linguae Sericae concept net.
- 20 July 2026 The Sino axis opens: Classical Chinese, the Chinese Δgamas, and the Man'yΕΕ‘Ε« β Phase 32 opens the Far Eastern axis: gold-annotated Classical Chinese, the Literary Chinese Δgamas beside their Pali parallels, the complete Man'yΕΕ‘Ε« lemmatized, and the reconstruction shelf from Old Chinese through the Qieyun to Japanese readings.
- 19 July 2026 v1.1.0: the library triples β Six phases promoted to a version: the Indic and Egyptian expansions, the Hebrew shelf, pre-Roman Italy and Sicily, and the Ancient Near East β 737,299 documents, 11.36 million passages, 69 sources, a 12.4-million-row gold lemma index in twenty-two languages.
- 19 July 2026 The Ancient Near East phase: Hittite at scale, the CDLI catalog, and a seven-legged Bible β Phase 31 brings the cuneiform world beyond ORACC: the TLHdig Hittite corpus, the CDLI's 353,156-artifact universal catalog, the eBL Fragmentarium, Ugaritic, a millennium of Syriac, the Sumerian literary canon in two editions β and the Peshitta as the alignment hub's seventh leg.
- 18 July 2026 The library since 1.0: a Slavic push, a Celtic axis, and canonical memory β Six phases after v1.0.0: Vasmer and the StarLing bases, the Slovenian historical dictionaries, the damaskini corpus, a whole Celtic axis with the first Old Irish gold lemmas, and a metadata framework of dossiers, notes, and a content census.
- 14 July 2026 Nabu 1.0 β Phase 19 closes with the library whole β every registered source live, canonical memory for one's own material β and Nabu cuts its first versioned release, v1.0.0.
- 14 July 2026 The library as of today β An inaugural stock-taking: 170,684 documents, 4.27 million passages, fifteen gold-lemma languages, and the tool families that read them.
- 14 July 2026 The machinery phase: quickstart, language cards, invariants β Phase 18 (PR #22): nabu quickstart, language cards for an 803-code universe, three comparativist adapters, and a postcondition checker.
- 13 July 2026 The sources phase: Coptic, inscriptions, Monier-Williams β Phase 17 (PR #21): four adapters in one day β Coptic Scriptorium, the Epigraphic Database Heidelberg, Monier-Williams, and four proto shelves.
- 13 July 2026 Fuzzy search and the links graph β Phase 16 (PR #20): trigram fuzzy search for damaged texts, and a persistent citation graph fed by parallels, formulas, and cognates.