🤖 AI Summary
This study investigates whether generative AI can alleviate administrative burdens associated with preparing for the UK’s Research Excellence Framework (REF). Method: Employing a viable systems model and reverse-engineering approach, ChatGPT was used to automatically score and rank 822 business and management research papers; critical thresholds aligned with REF quality ratings (1*–4*) were calibrated iteratively. A hybrid review framework—“AI-first screening followed by human verification”—was developed to enhance efficiency and consistency while preserving academic rigor. Contribution/Results: Empirical evaluation shows 49%–69% agreement between AI-generated scores and official REF outcomes across rating bands. The system effectively identifies borderline cases requiring prioritized human review, reducing manual assessment effort. This work constitutes the first integration of generative AI into REF’s internal quality assessment pipeline and introduces a verifiable, interpretable, and scalable hybrid evaluation paradigm.
📝 Abstract
This paper examines the potential for generative artificial intelligence (GenAI) to assist with internal review processes for research quality evaluations in UK higher education and particularly in preparation for the Research Excellence Framework (REF). Using the lens of function substitution in the Viable Systems Model, we present an experimental methodology using ChatGPT to score and rank business and management papers from REF 2021 submissions, "reverse engineering" the assessment by comparing AI-generated scores with known institutional results. Through rigourous testing of 822 papers across 11 institutions, we established scoring boundaries that aligned with reported REF outcomes: 49% between 1* and 2*, 59% between 2* and 3*, and 69% between 3* and 4*. The results demonstrate that AI can provide consistent evaluations that help identify borderline evaluation cases requiring additional human scrutiny while reducing the substantial resource burden of traditional internal review processes. We argue for application through a nuanced hybrid approach that maintains academic integrity while addressing the multi-million pound costs associated with research evaluation bureaucracy. While acknowledging these limitations including potential AI biases, the research presents a promising framework for more efficient, consistent evaluations that could transform current approaches to research assessment.