Layers

Every document in the library carries its text and its language. Three add-on layers answer the other questions a reader brings to it: when was this written, where does it come from, and what kind of text is it? The layers share one doctrine:

When — the dates layer

As of 19 September 2026 — a live census: 1,309,880 documents carry dating bounds (of 2,274,296 live documents, 57%). nabu search --century -21, nabu list --by-date and nabu vocab --by-century all read the same per-document interval.

Every dating bound is an honest year interval extracted from whatever dating form upstream actually ships:

What resists an honest parse stays raw and unbounded: prose datings (“Mitte 15. Jh.” without a grid), upstream’s own no-date fillers, circa strings with nothing firmer behind them. An interval is always a real range — no fake midpoints. Documents whose range spans several centuries are bucketed by their earliest bound below (496,308 such documents — the announced bias, not a hidden one).

The centuries

Century Documents Largest sources
85th c. BCE 218 cdli 218
35th c. BCE 418 cdli 418
34th c. BCE 1,872 cdli 1,861 · oracc 11
33rd c. BCE 186 cdli 186
32nd c. BCE 5,548 cdli 4,958 · oracc 590
31st c. BCE 1,730 cdli 1,729 · edr 1
29th c. BCE 1,886 cdli 1,200 · oracc 686
27th c. BCE 10,336 aes 10,333 · cdli 2 · elephantine 1
26th c. BCE 2,098 cdli 1,898 · oracc 200
25th c. BCE 5,720 cdli 5,086 · oracc 634
24th c. BCE 17,624 cdli 17,298 · oracc 309 · elephantine 15 · ebl 2
23rd c. BCE 4 elephantine 4
22nd c. BCE 6,136 cdli 5,331 · oracc 803 · elephantine 2
21st c. BCE 194,084 cdli 111,093 · oracc 81,610 · ebl 796 · aes 585
20th c. BCE 14,635 cdli 14,380 · oracc 249 · ebl 5 · elephantine 1
19th c. BCE 66,775 cdli 59,466 · oracc 6,346 · ebl 962 · elephantine 1
18th c. BCE 6,315 cdli 6,299 · ebl 14 · elephantine 2
17th c. BCE 1,415 cdli 1,413 · ebl 2
16th c. BCE 1,486 aes 1,484 · elephantine 1 · tla-hf 1
15th c. BCE 14,707 cdli 14,692 · oracc 15
14th c. BCE 23,196 cdli 17,485 · oracc 5,571 · ebl 139 · elephantine 1
13th c. BCE 843 cdli 841 · ebl 1 · elephantine 1
12th c. BCE 238 cdli 232 · elephantine 6
11th c. BCE 496 aes 491 · elephantine 4 · ebl 1
10th c. BCE 51,944 cdli 30,137 · ebl 13,707 · oracc 8,100
9th c. BCE 2,373 cdli 1,955 · oracc 418
8th c. BCE 6,817 oracc 4,106 · cdli 2,636 · ceipom 71 · elephantine 2 · +2 more
7th c. BCE 29,633 cdli 17,344 · ebl 6,852 · oracc 5,044 · ceipom 196 · +6 more
6th c. BCE 10,339 cdli 8,221 · oracc 580 · isicily 433 · edr 414 · +9 more
5th c. BCE 2,381 isicily 928 · edr 670 · elephantine 237 · ceipom 172 · +7 more
4th c. BCE 8,349 cdli 3,357 · oracc 1,437 · iip 1,281 · elephantine 768 · +10 more
3rd c. BCE 7,580 papyri-ddbdp 4,362 · ceipom 1,325 · edr 895 · elephantine 348 · +9 more
2nd c. BCE 7,029 papyri-ddbdp 3,414 · edr 1,156 · ceipom 1,061 · itant 337 · +9 more
1st c. BCE 16,058 edr 10,960 · edh 1,904 · papyri-ddbdp 1,731 · iip 629 · +9 more
1st c. CE 69,543 edr 40,748 · edh 19,295 · papyri-ddbdp 6,828 · isicily 931 · +6 more
2nd c. CE 67,803 edh 26,822 · edr 22,808 · papyri-ddbdp 15,934 · elephantine 1,382 · +6 more
3rd c. CE 24,513 papyri-ddbdp 8,444 · edh 7,486 · edr 7,256 · isicily 595 · +6 more
4th c. CE 14,451 papyri-ddbdp 5,811 · edr 3,507 · edh 3,193 · iip 663 · +5 more
5th c. CE 5,620 papyri-ddbdp 1,955 · edr 1,524 · edh 905 · iip 501 · +7 more
6th c. CE 6,077 papyri-ddbdp 4,373 · edr 572 · edh 476 · corpus-corporum 221 · +7 more
7th c. CE 5,170 papyri-ddbdp 4,345 · elephantine 225 · corpus-corporum 189 · edh 175 · +9 more
8th c. CE 13,706 rundata 11,555 · papyri-ddbdp 1,786 · corpus-corporum 179 · openiti 85 · +9 more
9th c. CE 2,012 openiti 612 · corpus-corporum 580 · elephantine 354 · rundata 329 · +10 more
10th c. CE 4,851 rundata 3,239 · openiti 921 · elephantine 311 · corpus-corporum 303 · +10 more
11th c. CE 11,801 rundata 10,158 · openiti 918 · corpus-corporum 579 · elephantine 74 · +9 more
12th c. CE 8,357 okhc 4,614 · rundata 2,071 · corpus-corporum 926 · openiti 593 · +10 more
13th c. CE 4,533 corpus-gysseling 2,191 · rundata 1,289 · openiti 805 · corpus-corporum 110 · +9 more
14th c. CE 1,634 openiti 877 · rundata 521 · okhc 65 · ref 41 · +9 more
15th c. CE 1,348 openiti 548 · okhc 230 · rundata 183 · sillok 107 · +10 more
16th c. CE 5,993 eebo-tcp 3,680 · okhc 1,325 · openiti 479 · sillok 137 · +11 more
17th c. CE 57,145 eebo-tcp 46,267 · okhc 8,974 · dta 1,124 · openiti 318 · +15 more
18th c. CE 130,153 okhc 128,264 · dta 1,019 · openiti 250 · bibyeonsa 140 · +13 more
19th c. CE 296,083 okhc 291,317 · dta 2,777 · ctilc 545 · imp 409 · +18 more
20th c. CE 57,770 okhc 55,193 · openiti 1,420 · dta 427 · ctilc 422 · +7 more
21st c. CE 848 openiti 848

Where — the places layer

nabu place Girsu answers with the gazetteer card and every source’s holdings at that place. The layer has two halves: the gazetteers, held locally, and the matching decisions that connect a source’s verbatim place-name to an identity.

The gazetteers. The library never queries a gazetteer online — it holds them as canonical assets with provenance and derives one namespaced place index:

Namespace What Held rows License
pleiades: Pleiades — THE ancient-world gazetteer 42,284 CC BY 3.0
tm: Trismegistos Geo — finest grain for Greco-Roman Egypt 64,857 CC BY-SA 4.0
cigs: CIGS — the cuneiform world’s site index 598 CC BY 4.0
np: nabu-places native records (minted by scholarship, evidence required) 0 — the lane is new CC BY 4.0

Namespaces are parallel claims; equivalences between them are crosswalk data with provenance (3,438 rows: CIGS’s own columns + a Wikidata harvest), never inferred.

The decisions registry. nabu-places records the matching judgments — which identity a source’s verbatim place-name string denotes — each reviewable, in the pattern of nabu-lects. Three real rows tell the story: CDLI’s "Girsu (mod. Tello)"matched cigs:GIR

Coverage. At the places program’s close (9 August 2026), 585,682 documents carried a machine place reference, of 708,905 that name a place at all — a 3.9× gain over the pre-program state; the largest sources sit at 74–87% matched (CDLI 250,484 of 337,572, Oracc 87%, EDR 86%, EDH’s 73,507 upstream-asserted). The maintained coverage detail lives in docs/places.md. Honest limits stay on the record: a dozen upstream references cite defective Pleiades ids (flagged loudly by the health invariants); long-tail names below the curated waves stay visibly unmatched until their wave lands.

What kind — the classification layer

As of 19 September 2026 — a live census: 1,953,835 documents carry a kind classification (85% of the kind-eligible library), across 21 class families. What kind of document is this — an epitaph, a receipt, a hymn, a school exercise? Each source ships its own genre vocabulary (EpiDoc inscription types, cuneiform catalogue genres, manuscript headings); those upstream labels fold onto one ruled cross-corpus class list (config/kind_classes.yml, with crosswalks to EAGLE, LCGFT and Getty AAT where those vocabularies recognize the class).

Classification is multi-label — a funerary poem is both — and the tree is optionally deep: a class is a family with named sub-classes (literary holds literary/poetry, literary/narrative, …), any prefix matches its whole family in queries, and a document upstream calls just “Literary” honestly carries the bare family head. The upstream claim is preserved verbatim beside the fold, so the ruled class never erases what the source actually said.

The classes

Class Documents Sources
historiography 1,203,274 14
administrative 321,411 9
funerary 114,758 9
legal 34,540 9
letter 32,950 7
literary 30,482 12
royal 29,882 3
dedicatory 22,857 6
scripture 19,659 14
lexical 15,254 3
ritual 13,528 4
mark 12,115 9
honorific 11,346 3
divination 10,447 5
building 8,016 4
scholarly 7,306 6
school 5,142 5
magic 3,270 6
boundary 2,924 2
hymn-prayer 2,680 6
exegesis 382 1

What stays honest

Two buckets are counted beside — never inside — the classes: unknown (46,773 documents) is upstream’s own “cannot determine”, a claim the library preserves rather than overwrites; unmapped (14,163 documents) holds values still awaiting a fold rule — a visible curation worklist, not a silent discard. Values reviewed and found not to be genre claims at all (a physical-layout tag, a copy marker) are declared as such in config and render as their own census section. Sources carrying nothing genre-shaped at all (80 sources, 320,461 documents) speak through their recorded postures, not through silence.

One query across the layers

nabu search --century -21                        # when
nabu place Girsu                                 # where: the card + holdings
nabu kind census                                 # what kind: the class board
nabu kind census --unmapped                      # the classification worklist
nabu search --kind literary/poetry --lang la     # a family narrowed to verse
nabu search lugal --place cigs:GIR --from -2200 --to -2000 --kind administrative

Every filter also rides the MCP tools (nabu_search, nabu_place, nabu_show) for conversational use, and every document card (nabu show) renders its date interval, place references and kind line side by side — the three layers on one card.