One million documents — the presses of England open the door

27 August 2026 · Nabu news

The library crossed its first million today: 1,014,312 documents — 89.9 million passages — in 146 languages. The wave that did it is the early-modern English print world.

EEBO-TCP Phase I: 25,368 texts, everything printed in England from Caxton’s era to 1700 that the Text Creation Partnership’s keyers transcribed by hand from the page images — no OCR anywhere in it. Milton and Bacon and Donne are here, and Raleigh’s farewell broadside (“EVen such is time, which takes in trust / Our youth, our age, and all we have”), but the mass of it is the period itself: sermons, petitions, almanacs, plague bills, and the pamphlet storm of the civil-war decades — two-thirds of the corpus is post-1640. Every file carries an explicit public-domain dedication, and the archive’s editors had already opened the door by grant besides.

The machinery mattered as much as the mass. The corpus rides the same parser family built last week for the Corpus of Middle English — proven here at eighty times the scale, with exactly one casualty in 25,368 texts: a book whose scribal suspension marks lost their letters in transcription, diagnosed to twenty orphaned macrons and recovered the same day. And because every text carries its imprint year, the registry’s new Early Modern English stage received 20,981 documents onto the timeline automatically — catalogue dates meeting period bands, no human sorting a single one.

The English descent ladder now stands populated at every rung: Old English verse and prose, the Middle English of Chaucer and the Gawain-poet, and the printed language of Shakespeare’s century — one search away from each other, and from the hundred and forty-odd languages beside them. Phase II — thirty-five thousand more — is one configuration line away, waiting for its moment.

← All news  ·  Atom feed