Buckets:

2.47 TB
2,999 files
Updated 18 days ago
Name
Size
manifest
raw
README.md2.64 kB
xet
README.md

Chronicling America — bulk OCR mirror (snapshot 2026-08-12)

A byte-for-byte copy of the bulk OCR archives published by the Library of Congress on the Chronicling America Datasets portal, as they stood on 2026-08-12.

Not affiliated with or endorsed by the Library of Congress. loc.gov remains the authoritative source.

Contents

Path What
raw/*.tar.bz2 2,997 batch archives, unmodified
manifest/2026-08-12-loc-bulk-ocr.json The portal's manifest at snapshot time — batch names, page counts, LCCNs, sizes, SHA-256s, ingest and archive-creation dates
logs/, state/ Bookkeeping from the copy process; not part of the data

2,997 archives · 2,470,556,510,435 bytes (2.47 TB) · 23,718,115 pages · 3,196,462 issues, across 53 NDNP awardee institutions.

Each archive expands to ALTO XML plus plain text, one pair per page — about 97% XML and 3% text by volume, roughly 8.2× the compressed size.

No page images. Images (JP2 and PDF, but not TIFF masters) are published separately by LoC as BagIt bags under chroniclingamerica.loc.gov/data/batches/.

Every archive matches the SHA-256 recorded in the manifest: 0 missing, 0 unexpected, 0 size mismatches, byte totals equal. To check one:

shasum -a 256 raw/<batch>.tar.bz2
# compare with that batch's "sha256" field in the manifest

Why the snapshot is dated

LoC is reprocessing the collection with NDNP-Open-OCR, replacing legacy OCR text with improved output. Reprocessed batches appear under an incremented version (_ver01_ver02) and the previous archive leaves the portal, so the published set changes over time. This copy records the state on 2026-08-12, when roughly 1.8% of the collection had been reprocessed. It is a dated snapshot, not a live copy.

Rights

The Library of Congress states that the newspapers in Chronicling America are in the public domain or have no known copyright restrictions. Titles published more than 95 years ago are in the public domain in their entirety; more recent titles are believed to be public domain but may contain copyrighted third-party material, and LoC is explicit that responsibility for an independent legal assessment rests with the person using the item.

No additional licence is asserted over this copy, and none could be — it is unmodified public-domain material. Please credit the Library of Congress and the National Digital Newspaper Program, a partnership between LoC and the National Endowment for the Humanities.

Total size
2.47 TB
Files
2,999
Last updated
Aug 17
Pre-warmed CDN
US EU US EU

Contributors