🤖 AI Summary
This study systematically evaluates the capability of large language models (LLMs) to detect sexist content on the EXIST 2024 Twitter dataset and, for the first time, introduces a multi-annotator demographic lens—e.g., age and gender—to analyze group-level biases in model outputs. Method: We conduct fine-grained detection experiments across multiple state-of-the-art LLMs and employ statistical modeling to quantify how demographic attributes influence inter-annotator agreement between models and human annotators. Results: While LLMs demonstrate baseline discriminatory text detection ability, they fail to replicate the diversity of human judgments on sexism—exhibiting systematic disparities across age and gender subgroups. Contribution: Our work exposes a fundamental limitation in current LLM fairness modeling: the absence of sociodemographically grounded calibration. We propose integrating pluralistic social perspectives into model development as a necessary pathway toward equitable AI, offering empirical evidence and methodological guidance for building more representative, ethically robust AI systems.
📝 Abstract
The use of Large Language Models (LLMs) has proven to be a tool that could help in the automatic detection of sexism. Previous studies have shown that these models contain biases that do not accurately reflect reality, especially for minority groups. Despite various efforts to improve the detection of sexist content, this task remains a significant challenge due to its subjective nature and the biases present in automated models. We explore the capabilities of different LLMs to detect sexism in social media text using the EXIST 2024 tweet dataset. It includes annotations from six distinct profiles for each tweet, allowing us to evaluate to what extent LLMs can mimic these groups' perceptions in sexism detection. Additionally, we analyze the demographic biases present in the models and conduct a statistical analysis to identify which demographic characteristics (age, gender) contribute most effectively to this task. Our results show that, while LLMs can to some extent detect sexism when considering the overall opinion of populations, they do not accurately replicate the diversity of perceptions among different demographic groups. This highlights the need for better-calibrated models that account for the diversity of perspectives across different populations.