PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对AI辅助系统评价报告不一致的问题,通过分析SciLitBench数据集,提出PRISMA-LLM框架以改进方法、评估和限制报告。
📝 Abstract
Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
Problem

Research questions and friction points this paper is trying to address.

Large language models
AI-assisted systematic reviews
reporting inconsistency
workflow audit
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

PRISMA-LLM
systematic reviews
large language models
reporting framework
workflow transparency
💼 Related Jobs
No related jobs found.
M
Miguel Zabaleta
Department of AI and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA
Baihan Lin
Baihan Lin
Tenure-Track Professor, Mount Sinai, Harvard University
Speech / NLPML / RL / BanditsComputational PsychiatryTheoretical NeuroscienceBio-Inspired AI