Does ChatGPT score research quality differently by gender?
This study presents the first large-scale empirical investigation into whether large language models (LLMs), specifically ChatGPT, exhibit gender-based scoring bias when evaluating research quality under anonymized conditions. Leveraging 89,744 journal articles submitted to the UK’s Research Excellence Framework (REF) 2021, the authors removed author identifiers and obtained ChatGPT-assigned scores, which were then compared against official REF ratings and textual complexity metrics. Results indicate that papers with male first authors received slightly higher ChatGPT scores in most disciplines—particularly in health, science, and engineering—and this gender gap was more pronounced than in the official REF assessments. No significant disparity emerged in single-authored humanities and social sciences research. The findings suggest that the observed bias likely stems from indirect factors such as research field or topic rather than writing style, underscoring the need for caution regarding latent biases when deploying LLMs in research evaluation.