Institution profile

Common Crawl Foundation

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

Aug 17, 2026

This study addresses the underestimation of truncation loss in web-PDF corpus evaluation when relying solely on document-level metrics. For the first time, we quantify the statistical discrepancy between document counts and token distributions. Through large-scale parsing, Gini coefficient analysis, and cross-corpus validation, we demonstrate that merely 3% of documents contain half of all tokens, while truncation results in 55–62% text loss with existing recovery tools achieving less than 15% restoration. These findings confirm severe distortion in single-unit document statistics. Consequently, this work proposes a dual-unit reporting standard incorporating both document and token metrics, which significantly enhances the accuracy and scientific rigor of corpus evaluation, establishing a new benchmark for PDF data processing.

0 citationsRead paper

Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls

Jul 15, 2026

This study addresses a key limitation in traditional web crawling analysis, which typically assumes URLs are uniformly distributed and thereby overlooks the heterogeneous “persistent core–dynamic periphery” structure inherent in real-world web graphs. To overcome this assumption, the authors propose a novel two-component urn model grounded in discovery curves and sliding windows, enabling the first quantitative estimation of both the core proportion and the dynamic evolution parameters of the periphery within a web corpus. By jointly fitting coverage and survival rates, the model demonstrates strong empirical validity on both Common Crawl and German academic web datasets. Results reveal a pronounced core–periphery organization in both collections, with further heterogeneity observed even within the dynamic periphery itself.

0 citationsRead paper

Colour Contrast on the Web: A WCAG 2.1 Level AA Compliance Audit of Common Crawl's Top 500 Domains

Feb 27, 2026

This study addresses the critical issue of insufficient color contrast on web pages, which severely compromises accessibility for users with visual impairments. Leveraging Common Crawl’s WARC archives, the authors conduct the first large-scale, server-load-free static analysis of CSS from the homepages of the top 500 domains, automatically computing foreground–background color contrast ratios against the WCAG 2.1 AA compliance threshold of 4.5:1. The analysis reveals that 40.9% of all examined color pairs fail to meet this standard, with a median site-level compliance rate of only 62.7%. Notably, merely 20.4% of sites achieve full compliance, underscoring that inadequate color contrast remains a pervasive accessibility flaw across mainstream websites. This work establishes a novel, reproducible paradigm for large-scale web accessibility auditing through archival data.

0 citationsRead paper
Recent publications

Latest Papers

Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics

Aug 17, 2026

This study addresses the underestimation of truncation loss in web-PDF corpus evaluation when relying solely on document-level metrics. For the first time, we quantify the statistical discrepancy between document counts and token distributions. Through large-scale parsing, Gini coefficient analysis, and cross-corpus validation, we demonstrate that merely 3% of documents contain half of all tokens, while truncation results in 55–62% text loss with existing recovery tools achieving less than 15% restoration. These findings confirm severe distortion in single-unit document statistics. Consequently, this work proposes a dual-unit reporting standard incorporating both document and token metrics, which significantly enhances the accuracy and scientific rigor of corpus evaluation, establishing a new benchmark for PDF data processing.

0 citationsRead paper

Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls

Jul 15, 2026

This study addresses a key limitation in traditional web crawling analysis, which typically assumes URLs are uniformly distributed and thereby overlooks the heterogeneous “persistent core–dynamic periphery” structure inherent in real-world web graphs. To overcome this assumption, the authors propose a novel two-component urn model grounded in discovery curves and sliding windows, enabling the first quantitative estimation of both the core proportion and the dynamic evolution parameters of the periphery within a web corpus. By jointly fitting coverage and survival rates, the model demonstrates strong empirical validity on both Common Crawl and German academic web datasets. Results reveal a pronounced core–periphery organization in both collections, with further heterogeneity observed even within the dynamic periphery itself.

0 citationsRead paper

Colour Contrast on the Web: A WCAG 2.1 Level AA Compliance Audit of Common Crawl's Top 500 Domains

Feb 27, 2026

This study addresses the critical issue of insufficient color contrast on web pages, which severely compromises accessibility for users with visual impairments. Leveraging Common Crawl’s WARC archives, the authors conduct the first large-scale, server-load-free static analysis of CSS from the homepages of the top 500 domains, automatically computing foreground–background color contrast ratios against the WCAG 2.1 AA compliance threshold of 4.5:1. The analysis reveals that 40.9% of all examined color pairs fail to meet this standard, with a median site-level compliance rate of only 62.7%. Notably, merely 20.4% of sites achieve full compliance, underscoring that inadequate color contrast remains a pervasive accessibility flaw across mainstream websites. This work establishes a novel, reproducible paradigm for large-scale web accessibility auditing through archival data.

0 citationsRead paper