Beyond Human-Likeness: Mapping the Scientific Critique Profiles of LLMs and Human Reviewers

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过对比人类与大语言模型在ICLR 2025评审中的表现,使用五种理论框架分析了科学批评的功能差异,揭示了两者在批判性评论上的不同特点。
📝 Abstract
Large language models (LLMs) are increasingly discussed as tools for peer review, but their value is often assessed through human-likeness, perceived usefulness, or textual overlap with reviewer comments. This study shifts attention from whether LLMs resemble human reviewers to what functions of scientific critique they perform. Using ICLR 2025 peer-review data, we compare human reviews with LLM reviews generated under baseline and expert prompts. We operationalize scientific critique through two review acts, weakness critique and scientific questioning, and annotate point-level review text using five theory-guided frameworks: Anderson's knowledge types, Toulmin's argumentation model, Graesser's question depth, SOLO cognitive complexity, and Hattie's feedback functions. The results reveal a differentiated critique profile. Human reviews placed greater emphasis on scientific framing and revision guidance, more often identifying higher-order weaknesses and asking questions oriented toward improvement. LLM reviews showed higher rates of explanatory depth, integrative reasoning, and explicit argument structuring. Expert prompting did not make LLM critique uniformly more human-like; it partially narrowed some gaps but mainly amplified LLM-specific tendencies toward integration and formal argumentation. These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Peer Review
Scientific Critique
Human-Likeness
Review Functions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scientific Critique Profiles
Human-Likeness Evaluation
Expert Prompting
Integrative Reasoning
Argument Structuring
🔎 Similar Papers
No similar papers found.
Y
Yunhan Yang
School of Information, Journalism and Communication, The University of Sheffield, Sheffield, UK
Mike Thelwall
Mike Thelwall
School of Information, Journalism and Communication, The University of Sheffield
scientometricsaltmetricssentiment analysissocial mediaartificial intelligence
G
Guoxiu He
School of Economics and Management, East China Normal University, Shanghai, China