nabu-lects

Why another language classification? Prior art and the case for lects

Every text in a digital library carries a language tag. This document explains why the existing tagging systems — ISO 639, Glottolog, BCP 47, Wiktionary’s code space, and the conventions of working corpora — cannot, alone or together, answer a question that any library of ancient and medieval texts must answer constantly: which language or dialect, at what broad stage of its historical development, is this document written in? It reviews each system on its merits, shows where each falls short of that requirement, and motivates the design this repository implements: a small, scoped registry of lects.


1. The problem

Nabu is a research library holding roughly a million documents — some 68 million text passages — across more than 120 language codes: from Sumerian administrative tablets and Ugaritic ritual texts through Greek drama, the Latin of Cicero, Jerome, and twelfth-century chroniclers, to Classical Chinese, Tibetan canons, and the medieval lyric of Galicia. Every document carries a standard language code. And the codes, taken at face value, routinely tell less than the truth:

Occasionally the standards get it right by accident: Tibetan has three clean ISO codes (otb Old Tibetan, xct Classical Tibetan, bod modern Tibetan), and English, French, German, and Irish each received historical-stage codes. But that coverage is accidental — it exists where a scholarly community happened to lobby for a code, and nowhere else.

The requirement, stated once: an identifier that names unambiguously which language/dialect at what broad historical development stage a document or dictionary entry belongs to — with genealogy recoverable, reconstructions distinguished from attested varieties, registers and orthographies honestly recorded, and every code in actual use mappable onto it. The sections below measure the existing systems against that requirement.


2. Prior art

2.1 The ISO 639 family

ISO 639 is the bibliographic standard: part 1 (2-letter codes, ~180), part 2 (3-letter bibliographic codes, ~480, including collective codes), part 3 (comprehensive individual languages, ~7,900, maintained by SIL International), and part 5 (language families and groups: gem Germanic, sla Slavic, roa Romance, ine Indo-European…). Since ISO 639:2023 these parts are consolidated into a single standard.

ISO 639-3 classes every language as living, extinct, constructed, or historic, and the definition of “historic” is load-bearing: “A language is listed as historic when it is considered to be distinct from any modern languages that are descended from it: for instance, Old English and Middle English… the language have a literature that is treated distinctly by the scholarly community.” So ISO does recognize historical stages as codable languages — but only where someone filed a successful change request. The result is accidental coverage:

Who draws the boundaries? For part 2, library cataloguing tradition; for part-3 additions, whoever files a change request that SIL’s registrar accepts (the batch of ancient and historic codes largely came from the LINGUIST List — §2.3). There is no principled periodization authority: enm ends at 1500 because bibliographers said so; the Helsinki Corpus ends Middle English at 1500 via its own period M4 (1420–1500); and neither cites the other.

ISO 639-6 — the hierarchy that died. There was a direct attempt at the whole problem. ISO 639-6, published 2009-11-17, assigned alpha-4 codes for “comprehensive coverage of language variants” in a recursive hierarchy, researched by the registration authority GeoLang — genealogy and variant depth in one code space, exactly the shape a historical library wants. It was withdrawn on 2014-11-25, citing concerns about its usefulness and maintainability, and the lack of standardization across recursive hierarchies. Its life and death is the central cautionary tale for this repository: a global hierarchical registry of language variants proved unmaintainable. (The design lesson drawn here is that a scoped, single-maintainer registry — a hundred-odd records covering one library’s actual holdings — is precisely the case where the same idea works.) ISO 639-4:2010, for completeness, contains guidelines and principles only, no codes.

2.2 Glottolog

Glottolog (glottolog.org) catalogues languoids at three levels — family, language, dialect — each with a stable Glottocode, under expert-curated genealogical classification. It is the best genealogy in existence, and it handles the attested-ancestor problem with a characteristic device:

2.3 MultiTree and the LINGUIST List codes

The LINGUIST List defined the supplemental ISO-format codes for ancient and historic languages that later entered ISO 639-3, and minted local-use codes where ISO declined — gkm for Medieval (Byzantine) Greek is a LINGUIST List code, which Wiktionary adopted and ISO still lacks. MultiTree aggregated hypothesis trees of language relationships, including proto-language nodes. The infrastructure is aging — LL-MAP is defunct, MultiTree updates have stalled — though Wikidata still tracks the codes as property P1232. A quarry for precedent when new codes are needed; not a foundation to build on.

2.4 BCP 47 and the IANA Language Subtag Registry

BCP 47 (RFC 5646) defines the tag grammar the web runs on: language[-extlang][-script][-region][-variant…][-extension][-x-privateuse]. Three findings matter here.

(i) The registry really does hold historical-stage and orthography variants — verified against the current IANA registry:

Variant Description Prefix
vaidika / laukika / itihasa / bauddha Vedic / Classical / Epic / Buddhist Hybrid Sanskrit sa
emodeng Early Modern English (1500–1700) en
1606nict Late Middle French (to 1606) frm
1694acad Early Modern French fr
1901 / 1996 German orthography, traditional / 1996 reform de
petr1708 / luna1918 Russian Petrine / post-1917 orthography ru
polyton / monoton Greek polytonic / monotonic el
bohoric / dajnko / metelko Slovene historical alphabets sl
baku1926 / tarask / alalc97 Turkic unified Latin, Taraškievica, transliteration schemes various

So BCP 47 already models stage (sa-vaidika, en-emodeng) and orthography (ru-petr1708) as variant subtags on a base language — two of the exact axes needed beyond genealogy. But the registered set is anecdotal: Sanskrit received a full stage suite (one motivated registrant, 2010); Greek received nothing diachronic; Latin nothing at all.

(ii) Could private-use subtags (la-x-vulgar, grc-x-koine) carry the whole need? Syntactically yes — x- subtags are unrestricted and always well-formed. But they are semantically opaque by definition (RFC 5646 §2.2.7: meaning by private agreement only), invisible to every external consumer, carry no genealogy and no dates, and cannot be validated. Registering real variants is possible — a template to ietf-languages@iana.org, reviewed by the Language Subtag Reviewer, typically a matter of weeks per subtag — but that publishes one project’s periodization decisions to a global registry one negotiation at a time, and still provides no family walk. The extlang mechanism is a closed set for macrolanguage bridging (zh-, ar-, ms-…), unusable for stages.

(iii) What is genuinely worth adopting: the grammar discipline (ordered axes, lowercase, registry-validated well-formedness), the registered variant names as ready-made orthography vocabulary (petr1708, 1901, bohoric…), and the demonstration that language-shaped codes with a stage variant are readable and sortable by humans.

2.5 Wiktionary’s code space (and Wikidata lexemes)

Wiktionary maintains the de-facto richest stage vocabulary in existence — the code space any etymological aggregator ends up joining against. Three tiers:

The la-vul lesson comes from Wiktionary’s own guidelines (Wiktionary:About Vulgar Latin): attested Vulgar Latin words are entered in mainspace as Latin with a “Vulgar Latin” label, while unattested comparative-method forms live under Reconstruction:Latin/… — that is, even Wiktionary internally splits what the single code la-vul externally fuses. Any downstream system that adopts the code inherits the fusion.

Governance: the code tables change by editor consensus; codes are renamed, merged, promoted (as gkm was), and occasionally deleted. Excellent as an input to a mapping table; unacceptable as the identifier layer of a library that promises stable citations.

Wikidata lexemes, for completeness: a lexeme’s language may be any Wikidata item, so “Medieval Latin” or “Koine Greek” can serve as a lexeme language wherever an item exists, and forms can carry orthography qualifiers. But coverage for ancient lects is thin, there is no controlled stage taxonomy, and the identifiers are opaque Q-numbers. An interoperability target someday; not a base.

2.6 Corpus and philological practice

What the working corpora actually do is the strongest signal in this survey:

The pattern: everyone keeps the code coarse and hangs stage off metadata and dates. Nobody has promoted stage into the identifier itself — because nobody else has needed one identifier to disambiguate holdings spanning a hundred-plus code systems at once. A library that does faces the gap directly.


3. Why none suffices

Eight requirements, from §1’s problem statement:

(a) an unambiguous language×stage identifier · (b) genealogy recoverable · (c) proto, ancient, medieval, and modern varieties in one system · (d) registers and sociolects honestly distinguished from stages · (e) every code in actual use mappable onto it · (f) stable, sortable, human-readable identifiers · (g) date-band anchoring · (h) maintainable by a small team.

System a b c d e f g h
ISO 639-1/2/3/5 ✗ (lat, grc) ✗ (no proto) ✗ (lzh) — (it is the source) ~ (names only) ✓ (frozen)
ISO 639-6 ~ ~ ~ withdrawn
Glottolog ✗ (no Koine, no Medieval Latin) ✗ (no proto) ~ ~ (opaque codes) ~ (in names) ✓ (consume-only)
MultiTree / LINGUIST List ~ ~ ~ ~ ~ ✗ (moribund)
BCP 47, registered variants ~ (where they exist) ~ ✗ (global process)
BCP 47, x- composition ✓ syntactically ~ ~ ~ (opaque semantics)
Wiktionary codes ✗ (la-vul) ~ (module chains) ~ (conflated) ✗ (wiki churn) ✗ (not ours to govern)
Wikidata lexemes ~ ~ ~ ~ ~ ~
Corpus metadata pattern per corpus ~ ✗ (no identifiers at all)

Stated fairly, per system:

No row passes. The nearest thing to a pass is a composition: Wiktionary’s coverage (c, e) + Glottolog’s genealogy (b) + BCP 47’s syntax and orthography vocabulary (f) + the corpus world’s date-band semantics (g), under a scoped registry that a small team can actually maintain (a, d, h).


4. The conclusion: compose, under one roof

That composition is what this repository implements. The unit is the lect — the standard linguistics term for “any variety of a language, without commitment to its status” — and each lect identifier composes five axes:

lect-id  =  anchor [ ":" stage ] [ "/" variety ] [ "~" script ] [ "@" ortho ]

Worked examples: lat:med — Medieval Latin, at last distinct from Cicero’s lat:cla; grc:koi — Koine Greek, which no standard codes; zho/lit — Literary Chinese as what it is, a register of Chinese rather than a stage; roa:pro — Proto-Romance, anchored to the Romance family with mode: reconstructed, cleanly separated from the attested Vulgar Latin (lat/vul) that shares its Wiktionary code. A bare anchor remains legal and means the anchor’s whole span — honest coarseness where nothing finer is known. Where ISO already codes the stage (ang, fro, gmh, otb, xct…), the lect is the bare anchor, and the mapping is the identity — the common case costs nothing.

Each registry record carries its Glottocode (making Glottolog’s genealogy available as a crosswalk rather than a maintenance burden), its parent lect, its date band with a source note (bands are per-language scholarly decisions, recorded, never derived), and its mode. Existing codes map onto lects per collection, with a per-document escape hatch for corpora that span centuries under one tag.

The scope is deliberately what ISO 639-6’s was not: not the world’s languages, but the languages a real library actually holds — on the order of a hundred registry records, maintained where the holdings are, growing only when a holding forces a distinction.

The registry itself — the full identifier grammar, the schema, the stage vocabularies, and the mapping discipline — is specified in this repository’s README and schema documentation.


Sources