🤖 AI Summary
The feasibility and performance limits of large language models (LLMs) in replacing human experts for qualitative coding—a core task in systematic reviews—remain empirically unestablished.
Method: This study conducts the first empirical comparison of GPT-4 and Kimi on real-world systematic review data, evaluating inter-coder agreement (measured by Cohen’s κ) across varying data scales and question complexities, with human expert coding as the benchmark.
Contribution/Results: Both LLMs achieve near-human consistency (κ ≥ 0.85) on small-scale, structured coding tasks but exhibit substantial degradation (κ ≤ 0.50) on open-ended, high-complexity problems. The findings delineate an evidence-based applicability threshold for LLM-assisted evidence synthesis and propose a task-characteristic–driven framework for model selection and human–AI collaboration. This work provides methodological guidance and empirical validation for the trustworthy integration of LLMs into systematic review workflows.
📝 Abstract
This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that the performance of LLMs fluctuates by data volume and question complexity for systematic reviews.