LeWiDi-2025 at NLPerspectives: The Third Edition of the Learning with Disagreements Shared Task
This work addresses the challenge of modeling and evaluating AI systems’ capacity to capture human judgment variability—such as disagreement and subjectivity. Methodologically, we (1) extend the LeWiDi benchmark to four tasks (paraphrase identification, irony/sarcasm detection, natural language inference) with ordinal annotations and individual-perspective prediction; (2) introduce the first integration of soft-label learning and annotator modeling, moving beyond hard-classification paradigms; and (3) propose a multi-task training framework jointly optimizing distributional prediction, individual annotator modeling, and population-level judgment distribution learning. Contributions include two novel evaluation metrics that surpass conventional measures like cross-entropy, and comprehensive empirical analysis revealing strengths and limitations of existing approaches in modeling judgment variability. These advances significantly enhance LeWiDi’s utility and extensibility as a benchmark platform for controversy-aware AI.