Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

📅 2025-04-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
The feasibility and performance limits of large language models (LLMs) in replacing human experts for qualitative coding—a core task in systematic reviews—remain empirically unestablished. Method: This study conducts the first empirical comparison of GPT-4 and Kimi on real-world systematic review data, evaluating inter-coder agreement (measured by Cohen’s κ) across varying data scales and question complexities, with human expert coding as the benchmark. Contribution/Results: Both LLMs achieve near-human consistency (κ ≥ 0.85) on small-scale, structured coding tasks but exhibit substantial degradation (κ ≤ 0.50) on open-ended, high-complexity problems. The findings delineate an evidence-based applicability threshold for LLM-assisted evidence synthesis and propose a task-characteristic–driven framework for model selection and human–AI collaboration. This work provides methodological guidance and empirical validation for the trustworthy integration of LLMs into systematic review workflows.

Technology Category

Application Category

📝 Abstract
This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that the performance of LLMs fluctuates by data volume and question complexity for systematic reviews.
Problem

Research questions and friction points this paper is trying to address.

Evaluating GPT-4 and Kimi for systematic review automation
Comparing LLM-generated codes with human-coded systematic reviews
Assessing LLM performance variability with data and question complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using GPT-4 for systematic reviews
Comparing LLM-generated and human codes
Evaluating performance by data complexity
D
Dandan Chen Kaptur
Pearson
Y
Yue Huang
Measurement Incoporated
Xuejun Ryan Ji
Xuejun Ryan Ji
British Columbia College of Nurses and Midwives
Y
Yanhui Guo
University of Illinois Springfield
B
Bradley Kaptur
HSHS Saint John’s Hospital