NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM在评估论文新颖性方面的不足,提出NovGauge基准测试,通过任务、问题和方法三个维度细粒度诊断模型性能。
📝 Abstract
Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non-hallucinated correct-positive judgments, over 70% cite evidence fails to logically support the stated reason. The best-performing model, GPT-5.5, achieves 43-72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding.
Problem

Research questions and friction points this paper is trying to address.

Large language models
novelty assessment
benchmark
faithful evidence
diagnosis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-Grained Benchmark
Novelty Assessment
Cascading Diagnostic Pipeline
Evidence Grounding
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
G
Guoqiang Zhang
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
K
Kexin Tan
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China; Graduate School of Engineering, The University of Tokyo, Tokyo, Japan
Ming Zhang
Ming Zhang
Shanghai Jiao Tong University
medical imaging
Li Ju
Li Ju
Department of Information Technology, Uppsala University
Federated LearningDistributed OptimizationUncertainty QuantificationMultimodal Language Models
W
Wenqing Jing
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Z
Zhonghan Yue
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Jiayi Chen
Jiayi Chen
Phd student of CS, Fudan University
Large Language Model
S
Shiqiang Wu
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
S
Shaofan Liu
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Y
Yue Zhang
Atom Infinite Pte. Ltd., Singapore
Y
Yuankai Ying
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Y
Yang Shi
AtomInnoLab
T
Tao Gui
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Qi Zhang
Qi Zhang
Fudan University
SAGINsatellite routing
X
Xuanjing Huang
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China