NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

📅 2026-04-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of systematic evaluation benchmarks for assessing large language models’ ability to judge the novelty of academic papers—a critical gap that hinders their effective support in peer review. To bridge this gap, the authors introduce NovBench, the first large-scale benchmark comprising 1,684 NLP conference papers along with expert-written novelty assessments. They propose a four-dimensional evaluation framework encompassing relevance, correctness, coverage, and clarity. By integrating novelty claims from paper introductions with expert judgments, and leveraging prompt engineering, fine-tuning, and a hybrid evaluation strategy combining automatic and human assessment, their experiments reveal significant deficiencies in current models’ capacity to comprehend scientific novelty and adhere to instructions, underscoring the need for targeted model improvements.

Technology Category

Application Category

📝 Abstract
Novelty is a core requirement in academic publishing and a central focus of peer review, yet the growing volume of submissions has placed increasing pressure on human reviewers. While large language models (LLMs), including those fine-tuned on peer review data, have shown promise in generating review comments, the absence of a dedicated benchmark has limited systematic evaluation of their ability to assess research novelty. To address this gap, we introduce NovBench, the first large-scale benchmark designed to evaluate LLMs' capability to generate novelty evaluations in support of human peer review. NovBench comprises 1,684 paper-review pairs from a leading NLP conference, including novelty descriptions extracted from paper introductions and corresponding expert-written novelty evaluations. We focus on both sources because the introduction provides a standardized and explicit articulation of novelty claims, while expert-written novelty evaluations constitute one of the current gold standards of human judgment. Furthermore, we propose a four-dimensional evaluation framework (including Relevance, Correctness, Coverage, and Clarity) to assess the quality of LLM-generated novelty evaluations. Extensive experiments on both general and specialized LLMs under different prompting strategies reveal that current models exhibit limited understanding of scientific novelty, and that fine--tuned models often suffer from instruction-following deficiencies. These findings underscore the need for targeted fine-tuning strategies that jointly improve novelty comprehension and instruction adherence.
Problem

Research questions and friction points this paper is trying to address.

novelty assessment
large language models
peer review
academic publishing
evaluation benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

NovBench
novelty assessment
large language models
peer review
evaluation benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wenqing Wu
School of Economics and Management, Nanjing University of Science and Technology
Yi Zhao
Yi Zhao
Indiana University School of Medicine
mediation analysisdecomposition methodsneuroimaging
Yuzhuo Wang
Yuzhuo Wang
Harbin Intitute of Technology
knowledge graph
S
Siyou Li
School of Electronic Engineering and Computer Science, Queen Mary University of London
J
Juexi Shao
School of Electronic Engineering and Computer Science, Queen Mary University of London
Yunfei Long
Yunfei Long
Michigan State University
Computer VisionSensor Fusion
Chengzhi Zhang
Chengzhi Zhang
Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science