Buckets:
Chronicling America — bulk OCR mirror (snapshot 2026-08-12)
A byte-for-byte copy of the bulk OCR archives published by the Library of Congress on the Chronicling America Datasets portal, as they stood on 2026-08-12.
Not affiliated with or endorsed by the Library of Congress. loc.gov remains the authoritative source.
Contents
| Path | What |
|---|---|
raw/*.tar.bz2 |
2,997 batch archives, unmodified |
manifest/2026-08-12-loc-bulk-ocr.json |
The portal's manifest at snapshot time — batch names, page counts, LCCNs, sizes, SHA-256s, ingest and archive-creation dates |
logs/, state/ |
Bookkeeping from the copy process; not part of the data |
2,997 archives · 2,470,556,510,435 bytes (2.47 TB) · 23,718,115 pages · 3,196,462 issues, across 53 NDNP awardee institutions.
Each archive expands to ALTO XML plus plain text, one pair per page — about 97% XML and 3% text by volume, roughly 8.2× the compressed size.
No page images. Images (JP2 and PDF, but not TIFF masters) are published
separately by LoC as BagIt bags under chroniclingamerica.loc.gov/data/batches/.
Every archive matches the SHA-256 recorded in the manifest: 0 missing, 0 unexpected, 0 size mismatches, byte totals equal. To check one:
shasum -a 256 raw/<batch>.tar.bz2
# compare with that batch's "sha256" field in the manifest
Why the snapshot is dated
LoC is reprocessing the collection with
NDNP-Open-OCR, replacing
legacy OCR text with improved output. Reprocessed batches appear under an
incremented version (_ver01 → _ver02) and the previous archive leaves the
portal, so the published set changes over time. This copy records the state on
2026-08-12, when roughly 1.8% of the collection had been reprocessed. It is a
dated snapshot, not a live copy.
Rights
The Library of Congress states that the newspapers in Chronicling America are in the public domain or have no known copyright restrictions. Titles published more than 95 years ago are in the public domain in their entirety; more recent titles are believed to be public domain but may contain copyrighted third-party material, and LoC is explicit that responsibility for an independent legal assessment rests with the person using the item.
No additional licence is asserted over this copy, and none could be — it is unmodified public-domain material. Please credit the Library of Congress and the National Digital Newspaper Program, a partnership between LoC and the National Endowment for the Humanities.
- Total size
- 2.47 TB
- Files
- 2,999
- Last updated
- Aug 17
- Pre-warmed CDN
- US EU US EU