Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain

📅 2026-06-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of majority voting in annotating hate and offensive speech, which obscures substantial annotator disagreements on subjective boundaries and leads models to treat contested judgments as objective truths. Focusing on the HateXplain dataset, the authors systematically analyze disagreement patterns and evaluate three modeling approaches: hard-label BERT, soft-label models, and per-annotator multi-head architectures. Through chi-square tests and confidence analysis, they demonstrate that all models suffer a 22–28 percentage point accuracy drop on contentious samples, with the multi-head model achieving only 0.245 accuracy on offense-related disagreements. Critically, standard evaluation metrics fail to capture this degradation, as models often exhibit spuriously high confidence in incorrect predictions. The work exposes the structural bias inherent in majority voting for sensitive content annotation and its detrimental impact on model evaluation reliability.
📝 Abstract
Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this aggregation is not neutral: 42.6% of all annotator disagreement in HateXplain concentrates specifically at the hate/offensive boundary, a pattern consistent with annotators applying different thresholds for where hate begins (chi-squared = 135.199, df = 2, p < 0.0001). Both a hard-label BERT model (Model A) and a soft-label model (Model B) drop 22 percentage points in accuracy from agreed posts (~80%) to disagreement posts (~58%), confirmed at p < 0.0001. A per-annotator multi-head model (Model C) widens this gap further to 28 points while collapsing offensive disagreement accuracy to 0.245. Critically, Model A expresses significantly higher confidence on boundary case errors than Model C (0.710 vs. 0.495, p < 0.0001), meaning standard evaluation metrics will not detect the failure. Three downstream interventions of increasing sophistication all fail to recover boundary accuracy. We argue the problem is structural. Majority vote presents a contested judgment as ground truth, and models inherit that false certainty. The intervention must be upstream in annotation design.
Problem

Research questions and friction points this paper is trying to address.

hate speech
annotator disagreement
majority vote
annotation bias
offensive boundary
Innovation

Methods, ideas, or system contributions that make the work stand out.

majority vote bias
annotator disagreement
hate/offensive boundary
annotation design
model confidence calibration
J
Joshua Muhumuza
Makerere University, Kampala, Uganda
J
Joab Ezra Agaba
Makerere University, Kampala, Uganda
M
Mercy Amiyo
Makerere University, Kampala, Uganda