Does ChatGPT score research quality differently by gender?

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study presents the first large-scale empirical investigation into whether large language models (LLMs), specifically ChatGPT, exhibit gender-based scoring bias when evaluating research quality under anonymized conditions. Leveraging 89,744 journal articles submitted to the UK’s Research Excellence Framework (REF) 2021, the authors removed author identifiers and obtained ChatGPT-assigned scores, which were then compared against official REF ratings and textual complexity metrics. Results indicate that papers with male first authors received slightly higher ChatGPT scores in most disciplines—particularly in health, science, and engineering—and this gender gap was more pronounced than in the official REF assessments. No significant disparity emerged in single-authored humanities and social sciences research. The findings suggest that the observed bias likely stems from indirect factors such as research field or topic rather than writing style, underscoring the need for caution regarding latent biases when deploying LLMs in research evaluation.
📝 Abstract
Large Language Models (LLMs) are being considered for research evaluation, raising concerns about the introduction of AI bias. This study investigates whether ChatGPT research quality scores differ by first-author gender using 89,744 journal articles from the UK Research Excellence Framework (REF) 2021. Author information was withheld from ChatGPT to avoid direct gender bias. Nevertheless, male first-authored papers had slightly higher ChatGPT scores in most Units of Assessment (UoAs), especially in health, science and engineering-related subjects, and this pattern was often stronger for ChatGPT than for REF scores, based on a departmental-level proxy. Rank-based ChatGPT gains relative to REF scores were also more favourable for male first-authored papers in most UoAs, although the differences were generally small. Gender differences were not evident for solo research in the social sciences, arts and humanities, however. The male-favouring pattern for first-authored research was not explained by gender differences in writing styles, at least as reflected in abstract complexity. Some ChatGPT-REF differences may also reflect the departmental averaging process used to generate the REF proxy scores. Average ChatGPT scores may differ by first-author gender indirectly through other factors, such as field, topic, method, journal context or authorship structure. Thus, this is an additional reason to be cautious with AI-based research evaluation.
Problem

Research questions and friction points this paper is trying to address.

gender bias
research evaluation
ChatGPT
Large Language Models
scientific quality assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI bias
research evaluation
gender disparity
Large Language Models
ChatGPT
🔎 Similar Papers
No similar papers found.