The Southeast Asia desk — and the first Old Mon text anywhere
2 September 2026 · Nabu news
Where Sanskrit traveled east it seeded a family of written vernaculars: Old Khmer on the stelae of Angkor, Old Cham on the coast of Vietnam, Old Malay in Srivijaya, Old Javanese in the court poetry of Majapahit, Pyu and Old Burmese on the plain of Bagan. The Southeast Asia desk — the library’s twenty-fifth — opened today with all of them at once: six sources, ~3,000 documents, ten languages the library had never held.
The epigraphic core is the DHARMA project’s corpora (EFEO/ERC, CC BY and BY-SA): the Corpus des inscriptions khmères — the Cœdès K-numbers, 1,219 editions of pre-Angkorian and Angkorian epigraphy; the Campā corpus, whose Đông Yên Châu inscription is the oldest attested Austronesian text in existence; the Nusantara charters (Old Malay from Kedukan Bukit to the Laguna copperplate, Old Sundanese, and the sīma charters of Java); and the Corpus of Pyu Inscriptions. Beside them stand the Old Burmese stone inscriptions of Bagan (1,121 faces, Myanmar-script Unicode with a full transliteration lane riding every line) and the desk’s literary wing: the kakawin critical editions — Deśavarṇana, Sutasoma, the Pararaton — with each canto’s meter riding its verses, and the Old Javanese Wordnet as glossary.
The find of the phase hid in the Pyu corpus. Old Mon — the language of the Dvaravati world and the fourth leg of the Myazedi quadrilingual — had no digital corpus anywhere; the surveys recorded the gap as absolute. It turns out two Old Mon inscriptions ride the Corpus of Pyu Inscriptions: seventy-nine lines, the language’s first machine-readable text, now sitting in the same library as everything it touched.
The ingest was an argument for defensive parsing. These are working scholarly repositories, and the first syncs surfaced their reality: inscriptions whose line numbering restarts on every face of the stone, a witness’s page-turns threaded through critical apparatus, editors resuming an interrupted line, duplicate stanza numbers in flagship editions, and one template placeholder reading “languageb”. Every shape was either given its honest citation form or quarantined loudly and recorded — nothing silently dropped, nothing silently merged.
The library now holds 2,268,754 documents — 106.1 million passages — in 167 language codes across 25 desks.