Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决OCR文本处理中信息丢失和重复工作的问题,提出了一种名为Enriched Text的方法,通过保留元数据并进行多语言标注来优化大规模文本处理。
📝 Abstract
Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single 'complete' stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.
Problem

Research questions and friction points this paper is trying to address.

OCR text
denoising
deduplicating
metadata
Innovation

Methods, ideas, or system contributions that make the work stand out.

Enriched Text
metadata preservation
multilingual support
deduplication
🔎 Similar Papers
No similar papers found.
David Lowry-Duda
David Lowry-Duda
ICERM
M
Matteo Cargnelutti
Institutional Data Initiative, Harvard Law School Library
C
Catherine Brobston
Institutional Data Initiative, Harvard Law School Library
S
Salwa Ismail
Harvard Library
G
Greg Leppert
Institutional Data Initiative, Harvard Law School Library
Amanda Watson
Amanda Watson
Harvard Law School Library
Jonathan Zittrain
Jonathan Zittrain
George Bemis Prof. of Law, Prof. of Computer Science, and Prof. of Public Policy, Harvard University
internet architectureprivacypropertyspeechgovernance