Layers
Every document in the library carries its text and its language. Three add-on layers answer the other questions a reader brings to it: when was this written, where does it come from, and what kind of text is it? The layers share one doctrine:
- Extracted, never guessed — every claim is read from what the source actually ships (a date attribute, a findspot, a genre label), and what resists an honest parse stays visibly unclaimed.
- Honesty buckets are first-class — “undated”, “unmatched”, “unmapped” and upstream’s own “cannot determine” are counted answers, never silent gaps.
- Derived and rebuildable — all three layers regenerate from the canonical sources on every rebuild; the upstream claim is preserved verbatim beside every fold.
- Composable — each layer is a search filter, and they stack:
nabu search "dis manibus" --kind funerary --from -100 --to 100 --place pleiades:423025.
When — the dates layer
As of 19 September 2026 — a live census:
1,309,880 documents carry dating
bounds (of 2,274,296 live
documents, 57%). nabu search
--century -21, nabu list --by-date and nabu vocab --by-century
all read the same per-document interval.
Every dating bound is an honest year interval extracted from whatever dating form upstream actually ships:
- Structured claims — EpiDoc
origDateattributes, catalogue year columns, machine date attrs (the Korean chronicles’ per-entry1617-01-00dates), publication-year headers. - Banded labels — Assyriological period labels (“Neo-Assyrian”) through a ruled period table; century-half grids (“15,2” = 15th c., 2nd half); anno-mundi annal years era-converted with the ambiguity kept as a span.
- Author dating — work-composition years and author life bands (Patrologia Latina), floruit strings (“fl. 892”).
What resists an honest parse stays raw and unbounded: prose datings (“Mitte 15. Jh.” without a grid), upstream’s own no-date fillers, circa strings with nothing firmer behind them. An interval is always a real range — no fake midpoints. Documents whose range spans several centuries are bucketed by their earliest bound below (496,308 such documents — the announced bias, not a hidden one).
The centuries
| Century | Documents | Largest sources |
|---|---|---|
| 85th c. BCE | 218 | cdli 218 |
| 35th c. BCE | 418 | cdli 418 |
| 34th c. BCE | 1,872 | cdli 1,861 · oracc 11 |
| 33rd c. BCE | 186 | cdli 186 |
| 32nd c. BCE | 5,548 | cdli 4,958 · oracc 590 |
| 31st c. BCE | 1,730 | cdli 1,729 · edr 1 |
| 29th c. BCE | 1,886 | cdli 1,200 · oracc 686 |
| 27th c. BCE | 10,336 | aes 10,333 · cdli 2 · elephantine 1 |
| 26th c. BCE | 2,098 | cdli 1,898 · oracc 200 |
| 25th c. BCE | 5,720 | cdli 5,086 · oracc 634 |
| 24th c. BCE | 17,624 | cdli 17,298 · oracc 309 · elephantine 15 · ebl 2 |
| 23rd c. BCE | 4 | elephantine 4 |
| 22nd c. BCE | 6,136 | cdli 5,331 · oracc 803 · elephantine 2 |
| 21st c. BCE | 194,084 | cdli 111,093 · oracc 81,610 · ebl 796 · aes 585 |
| 20th c. BCE | 14,635 | cdli 14,380 · oracc 249 · ebl 5 · elephantine 1 |
| 19th c. BCE | 66,775 | cdli 59,466 · oracc 6,346 · ebl 962 · elephantine 1 |
| 18th c. BCE | 6,315 | cdli 6,299 · ebl 14 · elephantine 2 |
| 17th c. BCE | 1,415 | cdli 1,413 · ebl 2 |
| 16th c. BCE | 1,486 | aes 1,484 · elephantine 1 · tla-hf 1 |
| 15th c. BCE | 14,707 | cdli 14,692 · oracc 15 |
| 14th c. BCE | 23,196 | cdli 17,485 · oracc 5,571 · ebl 139 · elephantine 1 |
| 13th c. BCE | 843 | cdli 841 · ebl 1 · elephantine 1 |
| 12th c. BCE | 238 | cdli 232 · elephantine 6 |
| 11th c. BCE | 496 | aes 491 · elephantine 4 · ebl 1 |
| 10th c. BCE | 51,944 | cdli 30,137 · ebl 13,707 · oracc 8,100 |
| 9th c. BCE | 2,373 | cdli 1,955 · oracc 418 |
| 8th c. BCE | 6,817 | oracc 4,106 · cdli 2,636 · ceipom 71 · elephantine 2 · +2 more |
| 7th c. BCE | 29,633 | cdli 17,344 · ebl 6,852 · oracc 5,044 · ceipom 196 · +6 more |
| 6th c. BCE | 10,339 | cdli 8,221 · oracc 580 · isicily 433 · edr 414 · +9 more |
| 5th c. BCE | 2,381 | isicily 928 · edr 670 · elephantine 237 · ceipom 172 · +7 more |
| 4th c. BCE | 8,349 | cdli 3,357 · oracc 1,437 · iip 1,281 · elephantine 768 · +10 more |
| 3rd c. BCE | 7,580 | papyri-ddbdp 4,362 · ceipom 1,325 · edr 895 · elephantine 348 · +9 more |
| 2nd c. BCE | 7,029 | papyri-ddbdp 3,414 · edr 1,156 · ceipom 1,061 · itant 337 · +9 more |
| 1st c. BCE | 16,058 | edr 10,960 · edh 1,904 · papyri-ddbdp 1,731 · iip 629 · +9 more |
| 1st c. CE | 69,543 | edr 40,748 · edh 19,295 · papyri-ddbdp 6,828 · isicily 931 · +6 more |
| 2nd c. CE | 67,803 | edh 26,822 · edr 22,808 · papyri-ddbdp 15,934 · elephantine 1,382 · +6 more |
| 3rd c. CE | 24,513 | papyri-ddbdp 8,444 · edh 7,486 · edr 7,256 · isicily 595 · +6 more |
| 4th c. CE | 14,451 | papyri-ddbdp 5,811 · edr 3,507 · edh 3,193 · iip 663 · +5 more |
| 5th c. CE | 5,620 | papyri-ddbdp 1,955 · edr 1,524 · edh 905 · iip 501 · +7 more |
| 6th c. CE | 6,077 | papyri-ddbdp 4,373 · edr 572 · edh 476 · corpus-corporum 221 · +7 more |
| 7th c. CE | 5,170 | papyri-ddbdp 4,345 · elephantine 225 · corpus-corporum 189 · edh 175 · +9 more |
| 8th c. CE | 13,706 | rundata 11,555 · papyri-ddbdp 1,786 · corpus-corporum 179 · openiti 85 · +9 more |
| 9th c. CE | 2,012 | openiti 612 · corpus-corporum 580 · elephantine 354 · rundata 329 · +10 more |
| 10th c. CE | 4,851 | rundata 3,239 · openiti 921 · elephantine 311 · corpus-corporum 303 · +10 more |
| 11th c. CE | 11,801 | rundata 10,158 · openiti 918 · corpus-corporum 579 · elephantine 74 · +9 more |
| 12th c. CE | 8,357 | okhc 4,614 · rundata 2,071 · corpus-corporum 926 · openiti 593 · +10 more |
| 13th c. CE | 4,533 | corpus-gysseling 2,191 · rundata 1,289 · openiti 805 · corpus-corporum 110 · +9 more |
| 14th c. CE | 1,634 | openiti 877 · rundata 521 · okhc 65 · ref 41 · +9 more |
| 15th c. CE | 1,348 | openiti 548 · okhc 230 · rundata 183 · sillok 107 · +10 more |
| 16th c. CE | 5,993 | eebo-tcp 3,680 · okhc 1,325 · openiti 479 · sillok 137 · +11 more |
| 17th c. CE | 57,145 | eebo-tcp 46,267 · okhc 8,974 · dta 1,124 · openiti 318 · +15 more |
| 18th c. CE | 130,153 | okhc 128,264 · dta 1,019 · openiti 250 · bibyeonsa 140 · +13 more |
| 19th c. CE | 296,083 | okhc 291,317 · dta 2,777 · ctilc 545 · imp 409 · +18 more |
| 20th c. CE | 57,770 | okhc 55,193 · openiti 1,420 · dta 427 · ctilc 422 · +7 more |
| 21st c. CE | 848 | openiti 848 |
Where — the places layer
nabu place Girsu answers with the gazetteer card and every source’s
holdings at that place. The layer has two halves: the gazetteers, held
locally, and the matching decisions that connect a source’s verbatim
place-name to an identity.
The gazetteers. The library never queries a gazetteer online — it holds them as canonical assets with provenance and derives one namespaced place index:
| Namespace | What | Held rows | License |
|---|---|---|---|
pleiades: |
Pleiades — THE ancient-world gazetteer | 42,284 | CC BY 3.0 |
tm: |
Trismegistos Geo — finest grain for Greco-Roman Egypt | 64,857 | CC BY-SA 4.0 |
cigs: |
CIGS — the cuneiform world’s site index | 598 | CC BY 4.0 |
np: |
nabu-places native records (minted by scholarship, evidence required) | 0 — the lane is new | CC BY 4.0 |
Namespaces are parallel claims; equivalences between them are crosswalk data with provenance (3,438 rows: CIGS’s own columns + a Wikidata harvest), never inferred.
The decisions registry.
nabu-places records the
matching judgments — which identity a source’s verbatim place-name
string denotes — each reviewable, in the pattern of
nabu-lects. Three real rows
tell the story: CDLI’s "Girsu (mod. Tello)" → matched cigs:GIR
pleiades:912855; EDR’s"Mediolanum"— six Pleiades places carry that title, the row says the Insubrian Milan and names the five rejected homonyms;"Irisagrig (mod. uncertain)"→ unlocatable — the site is unidentified in reality, and that is an answer, not a failure. An unlisted name is honestly unmatched; adapter-asserted upstream references always win over registry mints — the registry only ever fills silence.
Coverage. At the places program’s close (9 August 2026), 585,682 documents carried a machine place reference, of 708,905 that name a place at all — a 3.9× gain over the pre-program state; the largest sources sit at 74–87% matched (CDLI 250,484 of 337,572, Oracc 87%, EDR 86%, EDH’s 73,507 upstream-asserted). The maintained coverage detail lives in docs/places.md. Honest limits stay on the record: a dozen upstream references cite defective Pleiades ids (flagged loudly by the health invariants); long-tail names below the curated waves stay visibly unmatched until their wave lands.
What kind — the classification layer
As of 19 September 2026 — a live census:
1,953,835 documents carry a kind
classification (85% of the kind-eligible
library), across 21 class families. What
kind of document is this — an epitaph, a receipt, a hymn, a school
exercise? Each source ships its own genre vocabulary (EpiDoc
inscription types, cuneiform catalogue genres, manuscript headings);
those upstream labels fold onto one ruled cross-corpus class list
(config/kind_classes.yml, with crosswalks to EAGLE, LCGFT and Getty
AAT where those vocabularies recognize the class).
Classification is multi-label — a funerary poem is both — and the
tree is optionally deep: a class is a family with named
sub-classes (literary holds literary/poetry,
literary/narrative, …), any prefix matches its whole family in
queries, and a document upstream calls just “Literary” honestly
carries the bare family head. The upstream claim is preserved verbatim
beside the fold, so the ruled class never erases what the source
actually said.
The classes
| Class | Documents | Sources |
|---|---|---|
| historiography | 1,203,274 | 14 |
| administrative | 321,411 | 9 |
| funerary | 114,758 | 9 |
| legal | 34,540 | 9 |
| letter | 32,950 | 7 |
| literary | 30,482 | 12 |
| royal | 29,882 | 3 |
| dedicatory | 22,857 | 6 |
| scripture | 19,659 | 14 |
| lexical | 15,254 | 3 |
| ritual | 13,528 | 4 |
| mark | 12,115 | 9 |
| honorific | 11,346 | 3 |
| divination | 10,447 | 5 |
| building | 8,016 | 4 |
| scholarly | 7,306 | 6 |
| school | 5,142 | 5 |
| magic | 3,270 | 6 |
| boundary | 2,924 | 2 |
| hymn-prayer | 2,680 | 6 |
| exegesis | 382 | 1 |
What stays honest
Two buckets are counted beside — never inside — the classes: unknown (46,773 documents) is upstream’s own “cannot determine”, a claim the library preserves rather than overwrites; unmapped (14,163 documents) holds values still awaiting a fold rule — a visible curation worklist, not a silent discard. Values reviewed and found not to be genre claims at all (a physical-layout tag, a copy marker) are declared as such in config and render as their own census section. Sources carrying nothing genre-shaped at all (80 sources, 320,461 documents) speak through their recorded postures, not through silence.
One query across the layers
nabu search --century -21 # when
nabu place Girsu # where: the card + holdings
nabu kind census # what kind: the class board
nabu kind census --unmapped # the classification worklist
nabu search --kind literary/poetry --lang la # a family narrowed to verse
nabu search lugal --place cigs:GIR --from -2200 --to -2000 --kind administrative
Every filter also rides the MCP tools (nabu_search, nabu_place,
nabu_show) for conversational use, and every document card
(nabu show) renders its date interval, place references and kind
line side by side — the three layers on one card.