RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入RAP基准,使用LLM预测AI/ML领域未来六个月的研究趋势,发现历史累积信息有助于提高预测准确性,并通过微调改善了预测性能。
📝 Abstract
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.
Problem

Research questions and friction points this paper is trying to address.

Research Attention Prediction
Large language models
arXiv corpus
Exponentially weighted moving average
Forecast accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Research Attention Prediction
Large Language Models
Evidence Acquisition Biases
Fine-tuning