Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics
This study addresses the underestimation of truncation loss in web-PDF corpus evaluation when relying solely on document-level metrics. For the first time, we quantify the statistical discrepancy between document counts and token distributions. Through large-scale parsing, Gini coefficient analysis, and cross-corpus validation, we demonstrate that merely 3% of documents contain half of all tokens, while truncation results in 55–62% text loss with existing recovery tools achieving less than 15% restoration. These findings confirm severe distortion in single-unit document statistics. Consequently, this work proposes a dual-unit reporting standard incorporating both document and token metrics, which significantly enhances the accuracy and scientific rigor of corpus evaluation, establishing a new benchmark for PDF data processing.