Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出Think-Probe-Respond方法,通过探查和调整大语言模型在评估研究想法新颖性时的隐含判断,解决其倾向于给出中等新颖性评分的问题。
📝 Abstract
Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent "medium novelty" bias.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Novelty Judgment
Systematic Bias
Medium Novel
Innovation

Methods, ideas, or system contributions that make the work stand out.

Think-Probe-Respond
novelty judgment
large language models
latent judgments
medium novelty bias