<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://arvicco.github.io/nabu/feed.xml" rel="self" type="application/atom+xml" /><link href="https://arvicco.github.io/nabu/" rel="alternate" type="text/html" /><updated>2026-09-19T10:01:22+00:00</updated><id>https://arvicco.github.io/nabu/feed.xml</id><title type="html">Nabu</title><subtitle>A personal, local, license-honest library of the ancient world — searchable, citable, and rebuildable from canonical sources.</subtitle><entry><title type="html">v1.6.0 — four axes, and a public data wave</title><link href="https://arvicco.github.io/nabu/news/2026/09/18/v1-6-0-four-axes-and-a-data-wave/" rel="alternate" type="text/html" title="v1.6.0 — four axes, and a public data wave" /><published>2026-09-18T07:00:00+00:00</published><updated>2026-09-18T07:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/18/v1-6-0-four-axes-and-a-data-wave</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/18/v1-6-0-four-axes-and-a-data-wave/"><![CDATA[<p>Version 1.6.0 closes the classification arc that began with a survey
of twenty-one prior classification schemes and ended, three phases
later, with a working fourth axis. Beside <em>when</em>, <em>where</em> and <em>in
what language variety</em>, every document in the library can now answer
<strong>what kind of text it is</strong> — across
<strong>1,953,835 classified documents</strong>
(85% of the kind-eligible
library, as of 19 September 2026),
21 class families, every family
attested. The tree is optionally deep — <code>literary</code> holds its poetry,
narrative, drama and their siblings; any prefix matches its whole
family — and every fold preserves the upstream label verbatim beside
it. The three older layers now share one
<a href="/nabu/layers/"><strong>Layers</strong></a> page: one doctrine
(extracted, never guessed; honesty buckets counted, never hidden),
three sections, filters that stack in one query.</p>

<p>The release also ships the machinery a living library needs when an
<em>upstream</em> cleans house: when PerseusDL deliberately retired
ninety-six superseded Cicero editions, the library’s withdrawal alarm
said so — loudly, correctly, and forever. A reviewed upstream
curation event can now be <strong>accepted</strong>: the alarm quiets to a dated
note and re-arms the moment shedding grows past the accepted level.
The health board is fully green for the first time in a week, with
nothing swept under a rug to get there.</p>

<p>And the derived layers are going public. The
<a href="https://github.com/arvicco/nabu-data">nabu-data</a> sister repository
is receiving its largest wave since it opened: the complete
<strong>kind-classifications</strong> table (the fourth axis, per-document, with
every upstream label in-band), the <strong>cuneiform sense glosses</strong>
sidecar (Wiktionary’s Sumerian, Akkadian and Hittite lanes, under
their own share-alike license beside the CC-BY sign table), and
re-derivations of the standing catalog-wide datasets — lect
assignments, document dates, places — from a catalog two million
documents richer than their last cut. Each dataset carries its full
derivation provenance, as always: the exact producing code version,
the input identities, and a recipe that makes the build repeatable
end to end.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[The release that completes the library's fourth document axis and sends the layers out into the world: classification across two million documents, the unified Layers page, a clean health board, and the largest nabu-data publication wave since the repository opened.]]></summary></entry><entry><title type="html">Layers, and the literary family</title><link href="https://arvicco.github.io/nabu/news/2026/09/18/layers-and-the-literary-family/" rel="alternate" type="text/html" title="Layers, and the literary family" /><published>2026-09-18T06:00:00+00:00</published><updated>2026-09-18T06:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/18/layers-and-the-literary-family</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/18/layers-and-the-literary-family/"><![CDATA[<p>Four days after the classification axis went live, its class list got
its first structural lesson. The cuneiform catalogs label thousands of
tablets simply <em>Literary</em> — literature, no further claim — and the
original 26-head list had homes for poetry, narrative, drama, essay,
diary and wisdom, but none for their parent. Now it does, the natural
way: <strong><code>literary</code> is a family</strong>, the six literature classes live
inside it (<code>literary/poetry</code>, <code>literary/narrative/epic</code>, …), and a
document upstream calls just “Literary” honestly carries the bare
family head. Any prefix matches its whole family in queries —
<code>--kind literary</code> finds all of it,
<code>--kind literary/narrative</code> narrows, and the same deepening works for
every class. The live board:
<strong>21 class families, every one
attested, 1,953,835 documents
classified</strong> (85% of the
kind-eligible library, as of 19 September 2026).</p>

<p>The same pass drained the curation worklist. A full audit of every
unmapped upstream label folded roughly a hundred more values —
Hittite catalog ranges resolved against the corpus’s own sub-corpus
tags, the papyri’s German documentary tail, stray prayers, curse
tablets and writing exercises that already had homes — and introduced
an honest new category: values <strong>declared not-genre</strong> (a
physical-layout tag, a bare copy marker), reviewed and folded to
nothing, with the reason on record. The worklist now opens by saying
what it is and what to do with a line, and what remains in it is a
quarter of what was there before, all of it genuinely undecided.</p>

<p>And the site learned the lesson a reader taught it: the three axis
pages — dates, places, kinds — told one story in three places. They
are now one <a href="/nabu/layers/"><strong>Layers</strong></a> page: the
shared doctrine first (every layer extracted, never guessed; honesty
buckets counted, never hidden; everything derived and rebuildable),
then <em>when</em>, <em>where</em> and <em>what kind</em> as sections that compose into
one query:</p>

<pre><code>nabu search lugal --place cigs:GIR --from -2200 --to -2000 --kind administrative
</code></pre>

<p>Old links redirect; nothing is lost but the repetition.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[The classification tree learns depth — literature becomes one family with poetry, narrative, drama and their siblings inside it — the unmapped worklist shrinks to a quarter of its size, and the site's three axis pages merge into one Layers page.]]></summary></entry><entry><title type="html">The classification axis: what kind of document is this?</title><link href="https://arvicco.github.io/nabu/news/2026/09/14/the-classification-axis/" rel="alternate" type="text/html" title="The classification axis: what kind of document is this?" /><published>2026-09-14T06:00:00+00:00</published><updated>2026-09-14T06:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/14/the-classification-axis</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/14/the-classification-axis/"><![CDATA[<p>A gravestone inscription is <em>sepulcralis</em> in one collection, <em>epitaph</em>
in another, <em>funerary</em> in a third, <em>Inscription funéraire</em> in a
fourth, and a <em>Leichenpredigt</em> — a printed funeral sermon — in a
fifth. Until this week, “show me all funerary texts across the
library” required knowing every collection’s private jargon. Now it is
one query:</p>

<pre><code>nabu search "dis manibus" --kind funerary --lang la
</code></pre>

<p>The library’s fourth document axis — <strong>kind</strong>, beside dates, places,
and lects — folds each source’s own genre vocabulary onto a ruled
cross-corpus class list: 21
plain families (funerary, dedicatory, administrative, legal, letter,
lexical, scripture, divination, poetry, historiography, …), each
carrying pointers to the international vocabularies that recognize it
(the EAGLE epigraphic types, Library of Congress genre/form terms,
Getty AAT). No external standard spans cuneiform tablets, papyri,
stone, scripture, and novels at once — the survey behind this axis
verified that — so the list is the library’s own, and every upstream
label is preserved verbatim beside the fold.</p>

<p><strong>Classification is multi-label.</strong> An ode to a ruler is <em>poetry</em> AND
<em>royal</em>; a cuneiform “Administrative Letter” is both of its words; a
verse epitaph carries <em>funerary/epitaph</em> and <em>poetry</em> together. And
it is two-grain: <code>--kind divination</code> finds the whole family,
<code>--kind divination/extispicy</code> narrows to the liver omens.</p>

<p>The numbers, from the live census
(19 September 2026): <strong>1,953,835 documents classified —
85% of the library</strong> — drawn
from upstream catalogue genres (cuneiform’s CDLI and Oracc, the
Heidelberg and Rome epigraphic databases), the papyri’s HGV text
types (<em>Quittung</em>, <em>Vertrag</em>, <em>Mumienetikett</em>), the sibu 四部
classes of the Chinese canon, Japan’s NDC decimal codes, Sefaria’s
category tree, and whole-collection declarations where a source IS
one thing — the Korean dynastic records, the Buddhist and biblical
canons. What no rule covers yet sits in an honest, counted
<em>unmapped</em> bucket; sources carrying nothing genre-shaped say so
through their postures. The full board lives on the new
<a href="/nabu/kinds/">Kinds page</a>, and <code>nabu kind census</code>
prints it in under a second.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[The library gains its fourth axis. Beside when, where, and in what language variety, every document can now answer what KIND of text it is — funerary, administrative, letter, divination, historiography — one ruled vocabulary folded over every upstream jargon, with the original label preserved verbatim beside the fold.]]></summary></entry><entry><title type="html">The instruments phase: the library learns places and people</title><link href="https://arvicco.github.io/nabu/news/2026/09/11/the-instruments-phase/" rel="alternate" type="text/html" title="The instruments phase: the library learns places and people" /><published>2026-09-11T12:00:00+00:00</published><updated>2026-09-11T12:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/11/the-instruments-phase</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/11/the-instruments-phase/"><![CDATA[<p>The last three phases moved fast — the German print library (the
Deutsches Textarchiv, 5,478 volumes and 729,090 passages of German
from 1473 onward), a Slavic deepening (Slovenian protestant prints,
the Franček dictionary crosswalk, Old Church Slavonic additions), and
a long-tail sweep that carried the shelf past <em>Beowulf</em> in facsimile
and Hafez in Persian. The census stands at <strong>2,274,289 documents /
106,877,441 passages across 156 sources and 167 language codes</strong>.</p>

<p>This phase gave the library something different: <strong>instruments</strong> —
reference machinery that mints no texts of its own, but makes the
texts already held answerable in new ways.</p>

<p><strong>The library now knows where China is.</strong> The fourth gazetteer joins
Pleiades, Trismegistos and the cuneiform site index:
<a href="https://doi.org/10.7910/DVN/H3OB28">CHGIS/TGAZ</a> (Harvard–Fudan,
CC0), <strong>81,292 historical Chinese placenames</strong> from 221 BCE to 1911,
each carrying its hanzi name, pinyin transcription, valid years and
coordinates. <code>nabu place chgis:hvd_167661</code> answers instantly — 大川,
a village-town, coordinates and all.</p>

<p><strong>And the first text-mining lane ran over the Chinese canon.</strong> With a
gazetteer in hand, the library scanned all 4.57 million passages of
the Kanripo corpus — texts that carry no geographic metadata at all —
for exact Han-character matches against those placenames:
<strong>3,771,151 candidate attestations across 10,055 distinct names</strong>,
each recorded as evidence, never as fact. Precision rules do real
work here (single characters never match, over-common words are
stop-listed by measurement), and the candidates wait for human
review before any document is actually marked as speaking of a
place. The <em>Yijing</em> matching 大川 — “the great river” of the hexagram
formulas — is exactly why the review step exists.</p>

<p><strong>Persons arrive as an instrument too.</strong> The <a href="https://github.com/cbdb-project/cbdb_sqlite">China Biographical
Database</a> (Harvard /
Academia Sinica / Peking University; CC BY-NC-SA 4.0, confirmed at
first sync) — <strong>some 658,000 persons of Chinese history, 7th–19th
century</strong> — is now held and checksum-verified: names, dates, offices,
kinship, place associations. What surface it grows into (a person
card? person references on documents?) is a decision the library
will take deliberately, the way the places program grew from
gazetteers.</p>

<p><strong>The Arabic shelf learned its own intertextuality.</strong> From the
<a href="https://kitab-project.org">KITAB project</a>’s text-reuse statistics
(CC BY-NC-SA 4.0), the library minted <strong>931,943 reuse edges</strong>
between pairs of held OpenITI works — which books quote, excerpt and
rework which, with aligned-passage counts and chronology flags on
every edge. Asking <code>nabu links</code> on a held Arabic work now answers
with its textual relatives across eleven centuries.</p>

<p><strong>And one small, satisfying surface:</strong> the dictionary card now says
when a word was first written down. <code>nabu define bába --lang sl</code>
ends the Pleteršnik entry with <em>“first attested: Primož Trubar,
Katekizem, 1550”</em> — the crosswalk between a 19th-century dictionary
and the 16th-century corpus it describes, rendered where you look
words up. Alongside it, four more sources joined the dated timeline,
and a batch of era-bound performance assumptions was re-measured
against the hundred-million-passage reality.</p>

<p>As always: everything runs on one machine, every number above is
measured from the live catalog, and the licenses ride each record —
the instruments’ non-commercial grants are enforced by the tooling,
not by promise.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[Three research instruments enter the library — a historical-China gazetteer, a 658,000-person biographical database, and the KITAB text-reuse graph — and the first text-mining lane runs: 3.77 million place-name candidates over the Chinese canon, plus a dictionary card that now says when a word was first written down.]]></summary></entry><entry><title type="html">The search phase: six dragons, a 51-gigabyte diet, and an attic</title><link href="https://arvicco.github.io/nabu/news/2026/09/03/the-search-phase/" rel="alternate" type="text/html" title="The search phase: six dragons, a 51-gigabyte diet, and an attic" /><published>2026-09-03T11:00:00+00:00</published><updated>2026-09-03T11:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/03/the-search-phase</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/03/the-search-phase/"><![CDATA[<p>The library’s newest phase was spent entirely on <strong>search</strong> — four
lanes, each closing a different honest gap.</p>

<p><strong>Classical Chinese, Japanese and Korean now search exactly.</strong> These
shelves — the Buddhist canons, the Chosŏn state records, the Japanese
library, some eighteen million passages — write without word
boundaries, which quietly degraded them to second-class citizens of a
word-based index. The fix is a dedicated <strong>character-pair index</strong>:
23,372,988 pair rows over fifteen sources, so a Han or kana query
matches exact character sequences with no guessed segmentation
anywhere. The six dragons of the <em>Yijing</em>’s first hexagram, asked as
<code>nabu search 六龍</code>, answer in two-thirds of a second from the <em>Book of
Han</em>’s music treatise and two Buddhist commentaries quoting 乘六龍以御天
— “riding the six dragons to drive across the sky” — three corpora,
one query, mid-rebuild.</p>

<p><strong>The full-text engine went on a diet.</strong> The index had been storing a
private shadow copy of every passage it indexed — an artifact of the
engine’s default design, 39 gigabytes of duplication the catalog
already held. Rebuilt contentless and compacted, the full-text store
fell from <strong>87 GB to 36 GB</strong> with search behavior byte-for-byte
unchanged. Fifty-one gigabytes returned to the shelf, nothing lost but
redundancy.</p>

<p><strong>The library grew an attic — and it’s searchable.</strong> This collection
never hard-deletes: when an upstream source withdraws a document or a
revision prunes a passage, the text is withdrawn, not erased. Now
<code>search --withdrawn</code> searches exactly that shadow collection, and
every hit says <em>why</em> it left the open shelves — upstream gone, or
revision-pruned — so a vanished reading can be found, cited, and
traced years later.</p>

<p><strong>And the meaning-search machinery is built.</strong> The phase’s largest
single piece is the semantic lane: <code>nabu embed</code> distills each passage
of the literary core (4.9 million passages across seventeen Greek,
Latin, biblical and treebank sources) into a numerical fingerprint of
what it says, and <code>search --similar</code> — on the command line and
through the MCP server — answers “where else does the library say
this?” with ranked, banded, honestly-labelled neighbors:
cross-edition witnesses, drifting quotations, paraphrases. The whole
pipeline runs on this machine; no text leaves the box. The first
store build is a scheduled overnight run — and the lane will get the
announcement it deserves, with live walk-throughs, once its vectors
actually exist. Nothing on this site is pasted from imagination.</p>

<p>The plain-language write-ups landed with the code:
<a href="https://github.com/arvicco/nabu/blob/main/docs/embed.md">embed.md</a>
on the semantic lane, and
<a href="https://github.com/arvicco/nabu/blob/main/docs/lemma-enrichment.md">lemma-enrichment.md</a>
on the silver-lemma campaigns whose tool stack was also folded into
one-command installs this phase. Meanwhile the
<a href="/nabu/tools/">Tools</a> and
<a href="/nabu/examples/">Examples</a> pages have been
reworked around the new capabilities — the Tools page now carries the
MCP server as its own section, and the Examples page gained the
sinologist and the Old English scholar among its walk-throughs.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[One phase, all of it about finding things: exact Han-character search over the 23-million-pair CJK index, a full-text engine rebuilt to half its size with answers unchanged, a search view for everything the library ever withdrew — and the semantic-search machinery built end to end, awaiting its first overnight build.]]></summary></entry><entry><title type="html">The Southeast Asia desk — and the first Old Mon text anywhere</title><link href="https://arvicco.github.io/nabu/news/2026/09/02/the-indic-cosmopolis-eastern-edge/" rel="alternate" type="text/html" title="The Southeast Asia desk — and the first Old Mon text anywhere" /><published>2026-09-02T08:00:00+00:00</published><updated>2026-09-02T08:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/02/the-indic-cosmopolis-eastern-edge</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/02/the-indic-cosmopolis-eastern-edge/"><![CDATA[<p>Where Sanskrit traveled east it seeded a family of written
vernaculars: Old Khmer on the stelae of Angkor, Old Cham on the
coast of Vietnam, Old Malay in Srivijaya, Old Javanese in the court
poetry of Majapahit, Pyu and Old Burmese on the plain of Bagan. The
<strong>Southeast Asia desk</strong> — the library’s twenty-fifth — opened today
with all of them at once: six sources, ~3,000 documents, ten
languages the library had never held.</p>

<p>The epigraphic core is the <strong>DHARMA project’s corpora</strong> (EFEO/ERC,
CC BY and BY-SA): the Corpus des inscriptions khmères — the Cœdès
K-numbers, 1,219 editions of pre-Angkorian and Angkorian epigraphy;
the Campā corpus, whose Đông Yên Châu inscription is the oldest
attested Austronesian text in existence; the Nusantara charters
(Old Malay from Kedukan Bukit to the Laguna copperplate, Old
Sundanese, and the sīma charters of Java); and the Corpus of Pyu
Inscriptions. Beside them stand the <strong>Old Burmese stone
inscriptions of Bagan</strong> (1,121 faces, Myanmar-script Unicode with
a full transliteration lane riding every line) and the desk’s
literary wing: the <strong>kakawin critical editions</strong> — Deśavarṇana,
Sutasoma, the Pararaton — with each canto’s meter riding its
verses, and the Old Javanese Wordnet as glossary.</p>

<p>The find of the phase hid in the Pyu corpus. Old Mon — the language
of the Dvaravati world and the fourth leg of the Myazedi
quadrilingual — had no digital corpus anywhere; the surveys
recorded the gap as absolute. It turns out two Old Mon inscriptions
ride the Corpus of Pyu Inscriptions: seventy-nine lines, the
language’s first machine-readable text, now sitting in the same
library as everything it touched.</p>

<p>The ingest was an argument for defensive parsing. These are working
scholarly repositories, and the first syncs surfaced their reality:
inscriptions whose line numbering restarts on every face of the
stone, a witness’s page-turns threaded through critical apparatus,
editors resuming an interrupted line, duplicate stanza numbers in
flagship editions, and one template placeholder reading
“languageb”. Every shape was either given its honest citation form
or quarantined loudly and recorded — nothing silently dropped,
nothing silently merged.</p>

<p>The library now holds <strong>2,268,754 documents — 106.1 million
passages — in 167 language codes across 25 desks.</strong></p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[The library's 25th desk opens whole in one phase: the DHARMA epigraphic corpora (Old Khmer, Campā, Nusantara, Pyu), the kakawin library, the Old Burmese inscriptions of Bagan, and the Old Javanese Wordnet — ten languages never held before, among them two inscriptions of Old Mon, a language with no digital corpus anywhere until now.]]></summary></entry><entry><title type="html">A million records of Korea — the library doubles in a day</title><link href="https://arvicco.github.io/nabu/news/2026/09/01/a-million-records-of-korea/" rel="alternate" type="text/html" title="A million records of Korea — the library doubles in a day" /><published>2026-09-01T11:30:00+00:00</published><updated>2026-09-01T11:30:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/09/01/a-million-records-of-korea</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/09/01/a-million-records-of-korea/"><![CDATA[<p>The Korean desk opened in August on the state record: the
Veritable Records of Joseon, the Goryeosa family, the state
council’s daily register. Today it widened to the whole written
world around the throne. The historical slice of the <strong>Open Korean
Historical Corpus</strong> (Song et al., KAIST; CC BY-NC 4.0, with
personal-research ingestion welcomed by the authors) brings
<strong>1,198,779 records — 3.5 million passages</strong>: the <strong>Samguk sagi</strong>
of 1145, the oldest surviving Korean history; the <strong>Ilseongnok</strong>,
the Records of Daily Reflections kept by the Kyujanggak; the
<strong>munjip mass</strong> of the Institute for the Translation of Korean
Classics — the collected works of the literati, a shelf the
library had held only in slivers; the Gaksadeungnok local-office
records; and the old-literature databases of three national
archives. The census: 1.12 million records of hanmun — Literary
Sinitic as Korea wrote it — beside 58 thousand of Korean, with
Japanese legation records and a handful of true Middle Korean.</p>

<p>The licensing is the quiet achievement. The corpus is one
non-commercial dataset, but its records are not one thing: each
carries its own copyright status, and the adapter reads it
per record — <strong>1,198,691 relabel Public Domain</strong>, 15 carry Korea’s
KOGL attribution license, and 73 no-derivatives records earn no
upgrade at all and keep the stricter class. A million-record
source where every card tells the truth about its own terms.</p>

<p>The ingest itself argued for defensive design. The first sync
quarantined 2,761 records rather than guess: 2,760 turned out to
be English and French documents in the Copyright Commission’s
collection — classified and recovered the same day — and one had
no body at all. And a silent schema drift (the deposit spells its
copyright field differently than the published sample) was caught
not by any test but by spot-checking stored cards against the
source, then pinned in fixtures with real deposit bytes.</p>

<p>The same phase re-grained three Menota manuscripts whose pages
carried unnumbered lines — Old Norse prayer books and legal
manuscripts now cited by folio and running line, thirty thousand
lines where four hundred page-blocks stood. The library now
holds <strong>2,265,769 documents — 106.0 million passages — in 155
language codes.</strong> A month ago it had not yet crossed one million.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[The historical slice of the Open Korean Historical Corpus lands: 1,198,779 records — the Samguk sagi of 1145, the Ilseongnok court diaries, the munjip mass of the Korean literati — each carrying its own license label. The library's document count more than doubles in a single source.]]></summary></entry><entry><title type="html">The Iguvine Tables — and a hundred million passages</title><link href="https://arvicco.github.io/nabu/news/2026/08/31/the-iguvine-tables/" rel="alternate" type="text/html" title="The Iguvine Tables — and a hundred million passages" /><published>2026-08-31T21:00:00+00:00</published><updated>2026-08-31T21:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/08/31/the-iguvine-tables</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/08/31/the-iguvine-tables/"><![CDATA[<p>Seven bronze tables, found in 1444 in a theater vault at Gubbio,
carry the liturgy of a priestly brotherhood in a language Rome
erased: <strong>Umbrian</strong>. They are the longest ritual text surviving
from ancient Italy in any language — Latin included — and the
<strong>TITUS Osco-Umbrian corpus</strong> (text entry J. Gippert and V.
Slunečko; the only complete digital edition of the Tables) now
stands in the library under a personal grant extending the Avestan
terms: <strong>386 documents, 1,697 inscription lines</strong> — the Tables
complete, the Oscan inscriptions of Pompeii, Bantia and
Pietrabbondante beside them, and one archaic-Latin legal
comparandum the edition carries for contrast.</p>

<p>Each line arrives twice over: the edition’s unified Latin
transliteration is the searchable text, while the original script —
native Italic written right-to-left, Latin capitals, or Greek
letters in the far south — rides every line as an annotation naming
its alphabet. The registry minted an Umbrian anchor the same day,
completing the Sabellic pair; the comparative desk’s
Classical-Latin equivalence keys (from CEIPoM) already reach into
the Tables, so <code>search --lemma precor</code> finds <em>pesnimu</em> — labeled as
scholarly equivalence, never counted as attestation.</p>

<p>Two quieter milestones close the week. The 2026-08-31 census reads
<strong>102.4 million passages</strong> across <strong>1,066,990 documents</strong> — the
hundred-million line crossed without ceremony somewhere between the
early-modern presses and the rhyme books. And the rebuild machinery
completed its economics arc: derivation stamps, a builder-scoped
digest, an owner-attested trust bridge, and an LFS-aware identity
for the one shelf whose materialized payloads had read as dirt
forever — so the library that once took sixteen hours to rebuild
now re-verifies an unchanged catalog in minutes, honestly.</p>

<p>Numbers as of 2026-08-31.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[The seven bronze tables of Iguvium — the longest ritual text to survive from pre-Roman Italy, and the corpus of the Umbrian language — arrive with the Oscan inscriptions in the TITUS edition, under an extended personal grant. The same week's census crosses one hundred million passages, and the rebuild machinery learns to trust its own stamps.]]></summary></entry><entry><title type="html">The poets of early Akkadian — SEAL joins by grant</title><link href="https://arvicco.github.io/nabu/news/2026/08/31/the-poets-of-early-akkadian/" rel="alternate" type="text/html" title="The poets of early Akkadian — SEAL joins by grant" /><published>2026-08-31T09:00:00+00:00</published><updated>2026-08-31T09:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/08/31/the-poets-of-early-akkadian</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/08/31/the-poets-of-early-akkadian/"><![CDATA[<p><strong>SEAL — Sources of Early Akkadian Literature</strong> (Michael P. Streck
and Nathan Wasserman, Jerusalem/Leipzig) is where Assyriology keeps
its poetry: the Old Babylonian Gilgamesh tablets, love lyrics,
incantations, hymns and laments of the third and early second
millennium BCE, each in an up-to-date scholarly transliteration
with the damage honestly bracketed. It is the edition of record for
this literature, and it now stands in the library — <strong>408
compositions, 14,953 lines</strong> — under the editors’ written grant for
personal research use: local only, never redistributed, every
displayed line carrying the project’s credit.</p>

<p>The ingestion respected the edition’s variety the hard way. A first
crawl parsed only the corner of the corpus the samples had shown;
the full census then taught the parser four page shapes the site
actually serves — tabular scores with witness sigla, labeled
paragraphs, free prose — and the steady state reads every text page
the archive publishes, with three broken-markup pages quarantined
by name rather than served with holes.</p>

<p>On the dictionary side, the <strong>Coptic lexicon gained its second
etymological witness</strong>: the KELLIA project’s Egyptian-etymologies
table now rides beside the ORAEC crosswalk — two independent
scholarly witnesses to each Coptic word’s descent through Demotic
and hieroglyphic Egyptian, merged where they agree (they never
disagreed: a zero-conflict census preceded the merge) and credited
separately where each speaks alone.</p>

<p>Numbers as of 2026-08-31.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[Sources of Early Akkadian Literature — the edition of record for Old Babylonian Gilgamesh and the oldest Akkadian poetry — enters the library under its editors' written personal-research grant: 408 compositions at line grain, damage brackets and all. The Coptic dictionary gains a second etymological witness the same week.]]></summary></entry><entry><title type="html">The rhyme books arrive — an Old Mandarin desk, and the eastern shelves widen</title><link href="https://arvicco.github.io/nabu/news/2026/08/30/the-rhyme-books-arrive/" rel="alternate" type="text/html" title="The rhyme books arrive — an Old Mandarin desk, and the eastern shelves widen" /><published>2026-08-30T12:00:00+00:00</published><updated>2026-08-30T12:00:00+00:00</updated><id>https://arvicco.github.io/nabu/news/2026/08/30/the-rhyme-books-arrive</id><content type="html" xml:base="https://arvicco.github.io/nabu/news/2026/08/30/the-rhyme-books-arrive/"><![CDATA[<p>The Chinese rhyme-book tradition is the densest phonological record
any premodern language possesses: dictionaries organized not by
meaning but by <em>sound</em>, each one a snapshot of how the language was
pronounced in its century. This wave brings the tradition into the
library as a connected series, verified live on 2026-08-30.</p>

<p>At its head stand the <strong>Guangyun</strong> (1008, 25,336 entries) and <strong>Wang
Renxu’s Qieyun</strong> manuscript (17,227 entries) — the Middle Chinese
standard the whole discipline reconstructs from — joined by Fujita
Takumi’s <strong>restoration of the Qieyun itself</strong> (11,158 entries), the
lost 601 CE parent of them all, recovered from the surviving
fragments and quotations. Then the record jumps six centuries: the
<strong>蒙古字韻</strong> (1308, 9,446 entries), the ‘Phags-pa-script rhyme book
that writes Old Mandarin in the Mongol empire’s universal alphabet,
and the <strong>中原音韻</strong> of 1324 (5,877 entries), the arch-source of Old
Mandarin phonology, compiled for opera singers who needed to know
what actually rhymed. The registry minted a proper Old Mandarin
stage for them — no ISO code exists for it, so the lect layer now
carries what the standards cannot.</p>

<p>Beside the rhyme books, three more doors: a <strong>Classical↔Modern
Chinese parallel corpus</strong> — 14,608 chapter documents, 1.94 million
passages, ~967k sentence pairs of the classical canon facing its
modern rendering, the largest parallel shelf in the library after
the biblical hub; the <strong>Monlam Tibetan lexicon</strong> (449,829 entries),
the largest licensed machine-readable Tibetan headword list, which
more than doubled the dictionary layer’s total; and <strong>DACON</strong>, the
first annotated corpus of Classical Newar — Nepal’s
Sanskrit-cosmopolis literature, gold-tagged — a small shelf (4
texts, 977 passages) that brought a whole language family into the
registry.</p>

<p>Numbers as of the 2026-08-31 census; the shelves entered the
catalog with the 2026-08-29/30 first syncs and were verified and
wired the same week.</p>]]></content><author><name></name></author><category term="news" /><summary type="html"><![CDATA[Seven centuries of Chinese phonology land as a connected series of rhyme books — the Guangyun and Wang Renxu's Qieyun, Fujita's restoration of the Qieyun itself, the 'Phags-pa 蒙古字韻 and the 中原音韻 of 1324 — beside a 1.9-million-passage Classical↔Modern parallel corpus, the Monlam Tibetan lexicon, and the first annotated Newar shelf.]]></summary></entry></feed>