Beyond Polarization: The Generative Constraint of Chain-of-Thought in Pointwise Reranking

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了Chain-of-Thought模型在点式重排序中的表现问题,通过强化学习等方法尝试改善,但发现其排名差距难以克服,揭示了当前方法的瓶颈。
📝 Abstract
In pointwise document reranking, Chain-of-Thought models typically underperform direct scoring models. While existing diagnostics attribute this to inferior classification, score polarization, or calibration breakdown, whether targeted training can bridge this gap remains unclear. Our empirical study first confirms that this gap is stable across scales up to 32B parameters, ruling out model and data capacity confounders. We then apply stress tests utilizing reinforcement learning, fine-grained supervision, and architectural decoupling to explicitly repair these deviations. Although these interventions improve classification accuracy and absolute scores, the relative ranking gap persists. These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution, revealing a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought
pointwise reranking
classification
score polarization
calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought
pointwise reranking
reinforcement learning
ranking signal resolution