Every text in a digital library carries a language tag. This document explains why the existing tagging systems — ISO 639, Glottolog, BCP 47, Wiktionary’s code space, and the conventions of working corpora — cannot, alone or together, answer a question that any library of ancient and medieval texts must answer constantly: which language or dialect, at what broad stage of its historical development, is this document written in? It reviews each system on its merits, shows where each falls short of that requirement, and motivates the design this repository implements: a small, scoped registry of lects.
Nabu is a research library holding roughly a million documents — some 68 million text passages — across more than 120 language codes: from Sumerian administrative tablets and Ugaritic ritual texts through Greek drama, the Latin of Cicero, Jerome, and twelfth-century chroniclers, to Classical Chinese, Tibetan canons, and the medieval lyric of Galicia. Every document carries a standard language code. And the codes, taken at face value, routinely tell less than the truth:
lat covers Plautus, Cicero, the Vulgate, the
monastic charters of the twelfth century, and Vatican encyclicals — there is no ISO
code for Old, Late, Medieval, or Neo-Latin. grc (“Ancient Greek, to 1453”) spans
Homer, Attic drama, the Septuagint, the New Testament, and the whole Byzantine
millennium; the 1453 boundary is the fall of Constantinople, a library-cataloguing
convention rather than a linguistic one.lzh “Literary Chinese” names a written
register — wenyan — used continuously for some two millennia alongside the spoken
language’s actual stages (och Old Chinese, ltc Late Middle Chinese sit beside it
in ISO as if they were siblings of the same kind). In a large library this is not a
corner case: Literary Chinese is over 13 million passages here, the second-largest
holding.la-vul “Vulgar Latin” is used both for attested substandard Latin (the Appendix
Probi’s scoldings — “speculum non speclum”) and for the reconstructed Proto-Romance
of the comparative method (the DÉRom dictionary’s asterisked étymons). These are
categorically different objects — one is evidence, the other is inference — and any
machinery that treats reconstructions specially (asterisk display, etymological
closure) breaks when a single code means both. The mirror image also occurs: the code
gmq-pro “Proto-Norse” is worn by real Elder Futhark inscriptions — attested
epigraphy filed under a reconstruction label.arc shelf can hold Imperial
Aramaic documents, Qumran Biblical Aramaic, and the Targums’ Jewish Literary Aramaic
— a millennium of distinct varieties (here ISO is actually finer than common
practice: it splits oar, arc, jpa, tmr, syc, myz). Rabbinic and Medieval
Hebrew have no code at all beside hbo/heb. Church Slavonic’s chu covers both
canonical Old Church Slavonic and the later Russian and Serbian recensions.roa-opt) is one
medieval lect with two modern descendants, Portuguese and Galician — the family tree
branches underneath a historical stage, which no flat code list can express.Occasionally the standards get it right by accident: Tibetan has three clean ISO codes
(otb Old Tibetan, xct Classical Tibetan, bod modern Tibetan), and English, French,
German, and Irish each received historical-stage codes. But that coverage is accidental
— it exists where a scholarly community happened to lobby for a code, and nowhere else.
The requirement, stated once: an identifier that names unambiguously which language/dialect at what broad historical development stage a document or dictionary entry belongs to — with genealogy recoverable, reconstructions distinguished from attested varieties, registers and orthographies honestly recorded, and every code in actual use mappable onto it. The sections below measure the existing systems against that requirement.
ISO 639 is the bibliographic standard: part 1 (2-letter codes, ~180), part 2 (3-letter
bibliographic codes, ~480, including collective codes), part 3 (comprehensive individual
languages, ~7,900, maintained by SIL International), and part 5 (language families and
groups: gem Germanic, sla Slavic, roa Romance, ine Indo-European…). Since
ISO 639:2023 these parts are consolidated into a single standard.
ISO 639-3 classes every language as living, extinct, constructed, or historic, and the definition of “historic” is load-bearing: “A language is listed as historic when it is considered to be distinct from any modern languages that are descended from it: for instance, Old English and Middle English… the language have a literature that is treated distinctly by the scholarly community.” So ISO does recognize historical stages as codable languages — but only where someone filed a successful change request. The result is accidental coverage:
ang Old English (ca. 450–1100), enm Middle English (1100–1500) —
plus an IANA variant subtag en-emodeng for Early Modern English (no ISO code).fro Old French (842–ca. 1400), frm Middle French (ca. 1400–1600).goh Old High German (ca. 750–1050), gmh Middle High German
(ca. 1050–1500); gml and dum on the Low German / Dutch side.sga Old Irish (to 900), mga Middle Irish (900–1200).osp Old Spanish (to ca. 1500). Greek: grc vs ell, split at
1453 — the reference names and date bands descend from Library of Congress / MARC
cataloguing practice, not from historical linguistics.och Old Chinese, ltc Late Middle Chinese, and lzh Literary
Chinese — a register coded as a sibling of two stages, the conflation embodied in
the standard itself.otb, xct, bod — the clean case.qaa–qtz exists for codes with no ISO assignment.Who draws the boundaries? For part 2, library cataloguing tradition; for part-3
additions, whoever files a change request that SIL’s registrar accepts (the batch of
ancient and historic codes largely came from the LINGUIST List — §2.3). There is no
principled periodization authority: enm ends at 1500 because bibliographers said so;
the Helsinki Corpus ends Middle English at 1500 via its own period M4 (1420–1500); and
neither cites the other.
ISO 639-6 — the hierarchy that died. There was a direct attempt at the whole problem. ISO 639-6, published 2009-11-17, assigned alpha-4 codes for “comprehensive coverage of language variants” in a recursive hierarchy, researched by the registration authority GeoLang — genealogy and variant depth in one code space, exactly the shape a historical library wants. It was withdrawn on 2014-11-25, citing concerns about its usefulness and maintainability, and the lack of standardization across recursive hierarchies. Its life and death is the central cautionary tale for this repository: a global hierarchical registry of language variants proved unmaintainable. (The design lesson drawn here is that a scoped, single-maintainer registry — a hundred-odd records covering one library’s actual holdings — is precisely the case where the same idea works.) ISO 639-4:2010, for completeness, contains guidelines and principles only, no codes.
Glottolog (glottolog.org) catalogues languoids at three levels — family, language, dialect — each with a stable Glottocode, under expert-curated genealogical classification. It is the best genealogy in existence, and it handles the attested-ancestor problem with a characteristic device:
lati1261, level language, extinct) sits at the bottom of a chain of
pseudo-family nodes: Indo-European > Classical Indo-European > Italic >
Latino-Faliscan > Latinic > … > Imperial Latin (impe1234, level: family) >
Latin. The Romance languages hang under that same Imperial Latin node — that is,
Glottolog expresses “Latin is the ancestor of Romance” by minting a family node named
after the ancestor and placing the attested ancestor inside it as a leaf sister
of its own descendants.olde1238, “Old English (ca. 450–1100)”, ISO ang) likewise:
… Anglic > Anglo-Saxon > Old English, with Middle English a separate languoid further
down the Anglic tree. Note that the timespan lives in the languoid’s name, not in a
data field.grc/ell, nothing for Medieval Latin.The LINGUIST List defined the supplemental ISO-format codes for ancient and historic
languages that later entered ISO 639-3, and minted local-use codes where ISO declined —
gkm for Medieval (Byzantine) Greek is a LINGUIST List code, which Wiktionary adopted
and ISO still lacks. MultiTree aggregated hypothesis trees of language relationships,
including proto-language nodes. The infrastructure is aging — LL-MAP is defunct,
MultiTree updates have stalled — though Wikidata still tracks the codes as property
P1232. A quarry for precedent when new codes are needed; not a foundation to build on.
BCP 47 (RFC 5646) defines the tag grammar the web runs on:
language[-extlang][-script][-region][-variant…][-extension][-x-privateuse]. Three
findings matter here.
(i) The registry really does hold historical-stage and orthography variants — verified against the current IANA registry:
| Variant | Description | Prefix |
|---|---|---|
vaidika / laukika / itihasa / bauddha |
Vedic / Classical / Epic / Buddhist Hybrid Sanskrit | sa |
emodeng |
Early Modern English (1500–1700) | en |
1606nict |
Late Middle French (to 1606) | frm |
1694acad |
Early Modern French | fr |
1901 / 1996 |
German orthography, traditional / 1996 reform | de |
petr1708 / luna1918 |
Russian Petrine / post-1917 orthography | ru |
polyton / monoton |
Greek polytonic / monotonic | el |
bohoric / dajnko / metelko |
Slovene historical alphabets | sl |
baku1926 / tarask / alalc97 … |
Turkic unified Latin, Taraškievica, transliteration schemes | various |
So BCP 47 already models stage (sa-vaidika, en-emodeng) and orthography
(ru-petr1708) as variant subtags on a base language — two of the exact axes needed
beyond genealogy. But the registered set is anecdotal: Sanskrit received a full stage
suite (one motivated registrant, 2010); Greek received nothing diachronic; Latin
nothing at all.
(ii) Could private-use subtags (la-x-vulgar, grc-x-koine) carry the whole need?
Syntactically yes — x- subtags are unrestricted and always well-formed. But they are
semantically opaque by definition (RFC 5646 §2.2.7: meaning by private agreement
only), invisible to every external consumer, carry no genealogy and no dates, and
cannot be validated. Registering real variants is possible — a template to
ietf-languages@iana.org, reviewed by the Language Subtag Reviewer, typically a matter
of weeks per subtag — but that publishes one project’s periodization decisions to a
global registry one negotiation at a time, and still provides no family walk. The
extlang mechanism is a closed set for macrolanguage bridging (zh-, ar-, ms-…),
unusable for stages.
(iii) What is genuinely worth adopting: the grammar discipline (ordered axes,
lowercase, registry-validated well-formedness), the registered variant names as
ready-made orthography vocabulary (petr1708, 1901, bohoric…), and the
demonstration that language-shaped codes with a stage variant are readable and
sortable by humans.
Wiktionary maintains the de-facto richest stage vocabulary in existence — the code space any etymological aggregator ends up joining against. Three tiers:
gkm (Medieval Greek, promoted from etymology-only to full language), roa-opt
(Old Galician-Portuguese), zle-ort (Old Ruthenian), and others.la-vul
Vulgar, la-lat Late, la-med Medieval, la-ecc Ecclesiastical, la-ren
Renaissance, la-new New, with Old Latin as itc-ola), grc-koi Koine Greek, and
dozens more.<family>-pro codes (ine-pro, gem-pro,
sla-pro…) whose entries live in a dedicated Reconstruction: namespace, plus
family codes reusing ISO 639-5. Parent/ancestor chains are machine-readable in the
project’s module data, so a genealogical walk exists — though as a wiki artifact.The la-vul lesson comes from Wiktionary’s own guidelines (Wiktionary:About
Vulgar Latin): attested Vulgar Latin words are entered in mainspace as Latin with a
“Vulgar Latin” label, while unattested comparative-method forms live under
Reconstruction:Latin/… — that is, even Wiktionary internally splits what the single
code la-vul externally fuses. Any downstream system that adopts the code inherits
the fusion.
Governance: the code tables change by editor consensus; codes are renamed, merged,
promoted (as gkm was), and occasionally deleted. Excellent as an input to a mapping
table; unacceptable as the identifier layer of a library that promises stable
citations.
Wikidata lexemes, for completeness: a lexeme’s language may be any Wikidata item, so “Medieval Latin” or “Koine Greek” can serve as a lexeme language wherever an item exists, and forms can carry orthography qualifiers. But coverage for ancient lects is thin, there is no controlled stage taxonomy, and the identifiers are opaque Q-numbers. An interoperability target someday; not a base.
What the working corpora actually do is the strongest signal in this survey:
grc, lat, got, chu, orv, ang) with stage carried by corpus membership
and per-text dates — TOROT’s orv deliberately spans Old East Slavic through
Middle Russian.la; UD_Ancient_Greek-PROIEL
mixes Herodotus and the New Testament under one grc. UD_Old_French and
UD_Classical_Chinese exist exactly where ISO happened to grant a code — the
conflations downstream projects inherit are UD’s.@xml:lang is BCP 47 by mandate; the Guidelines point to private-use and
variant subtags for finer distinctions and to <langUsage> prose for the rest —
TEI outsources the problem to BCP 47 and metadata.The pattern: everyone keeps the code coarse and hangs stage off metadata and dates. Nobody has promoted stage into the identifier itself — because nobody else has needed one identifier to disambiguate holdings spanning a hundred-plus code systems at once. A library that does faces the gap directly.
Eight requirements, from §1’s problem statement:
(a) an unambiguous language×stage identifier · (b) genealogy recoverable · (c) proto, ancient, medieval, and modern varieties in one system · (d) registers and sociolects honestly distinguished from stages · (e) every code in actual use mappable onto it · (f) stable, sortable, human-readable identifiers · (g) date-band anchoring · (h) maintainable by a small team.
| System | a | b | c | d | e | f | g | h |
|---|---|---|---|---|---|---|---|---|
| ISO 639-1/2/3/5 | ✗ (lat, grc) |
✗ | ✗ (no proto) | ✗ (lzh) |
— (it is the source) | ✓ | ~ (names only) | ✓ (frozen) |
| ISO 639-6 | ~ | ✓ | ~ | ~ | ✗ | ✗ | ✗ | ✗ withdrawn |
| Glottolog | ✗ (no Koine, no Medieval Latin) | ✓ | ✗ (no proto) | ✗ | ~ | ~ (opaque codes) | ~ (in names) | ✓ (consume-only) |
| MultiTree / LINGUIST List | ~ | ~ | ~ | ✗ | ~ | ~ | ✗ | ✗ (moribund) |
| BCP 47, registered variants | ~ (where they exist) | ✗ | ✗ | ✗ | ~ | ✓ | ✗ | ✗ (global process) |
BCP 47, x- composition |
✓ syntactically | ✗ | ~ | ~ | ✓ | ~ (opaque semantics) | ✗ | ✓ |
| Wiktionary codes | ✗ (la-vul) |
~ (module chains) | ✓ | ~ (conflated) | ✓ | ✗ (wiki churn) | ✗ | ✗ (not ours to govern) |
| Wikidata lexemes | ~ | ~ | ~ | ~ | ~ | ✗ | ~ | ✗ |
| Corpus metadata pattern | ✓ per corpus | ✗ | ✗ | ~ | ✓ | ✗ (no identifiers at all) | ✓ | ✓ |
Stated fairly, per system:
lzh) as if it were a stage. It is where identifiers should be anchored, not
where the distinctions live.lati1261) are not
human-readable. It earns its place as a crosswalk field, not as the identifier.la-vul
attested/reconstructed fusion this work needs to kill, under wiki governance where
codes are renamed and merged by consensus. Map from it; do not build on it.No row passes. The nearest thing to a pass is a composition: Wiktionary’s coverage (c, e) + Glottolog’s genealogy (b) + BCP 47’s syntax and orthography vocabulary (f) + the corpus world’s date-band semantics (g), under a scoped registry that a small team can actually maintain (a, d, h).
That composition is what this repository implements. The unit is the lect — the standard linguistics term for “any variety of a language, without commitment to its status” — and each lect identifier composes five axes:
lect-id = anchor [ ":" stage ] [ "/" variety ] [ "~" script ] [ "@" ortho ]
Worked examples: lat:med — Medieval Latin, at last distinct from Cicero’s lat:cla;
grc:koi — Koine Greek, which no standard codes; zho/lit — Literary Chinese as what
it is, a register of Chinese rather than a stage; roa:pro — Proto-Romance, anchored
to the Romance family with mode: reconstructed, cleanly separated from the attested
Vulgar Latin (lat/vul) that shares its Wiktionary code. A bare anchor remains legal
and means the anchor’s whole span — honest coarseness where nothing finer is known.
Where ISO already codes the stage (ang, fro, gmh, otb, xct…), the lect is
the bare anchor, and the mapping is the identity — the common case costs nothing.
Each registry record carries its Glottocode (making Glottolog’s genealogy available as a crosswalk rather than a maintenance burden), its parent lect, its date band with a source note (bands are per-language scholarly decisions, recorded, never derived), and its mode. Existing codes map onto lects per collection, with a per-document escape hatch for corpora that span centuries under one tag.
The scope is deliberately what ISO 639-6’s was not: not the world’s languages, but the languages a real library actually holds — on the order of a hundred registry records, maintained where the holdings are, growing only when a holding forces a distinction.
The registry itself — the full identifier grammar, the schema, the stage vocabularies, and the mapping discipline — is specified in this repository’s README and schema documentation.
lati1261, impe1234, olde1238