What Are They Filtering Out? A Survey of Filtering Strategies for Harm Reduction in Pretraining Datasets
Pretraining data filtering strategies intended to reduce harmful content inadvertently exacerbate representational underrepresentation of marginalized groups, thereby amplifying demographic bias at the data level. Method: We systematically reviewed 55 English-language large language model technical reports to construct the first integrated data governance evaluation framework balancing safety and fairness. Through controlled experiments and quantitative bias analysis across mainstream filtering strategies, we measured their impact on group-level representation. Contribution/Results: Our analysis reveals that such strategies reduce text associated with disadvantaged groups by 12.7%–38.4% on average—significantly worsening representational disparity. This study provides the first empirical evidence refuting the “safety implies fairness” assumption in AI governance. We propose a co-optimization paradigm that jointly addresses content safety and equitable group representation, advocating for fairness-aware data curation in foundation model development.