ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives
为解决长篇叙事材料分析难题,提出ClueWeaver框架,通过奖励引导的双代理机制实现证据检索与解释,优化紧凑型本地模型性能。
为解决长篇叙事材料分析难题,提出ClueWeaver框架,通过奖励引导的双代理机制实现证据检索与解释,优化紧凑型本地模型性能。
本文提出MulVec方法,通过细粒度角色感知匹配解决无需训练的零样本组合图像检索问题,提高检索精度。
Existing multi-hop fact verification methods often suffer from reasoning deviations or erroneous conclusions due to a lack of global objective awareness and conflicts between parametric knowledge and retrieved evidence. To address these issues, this work proposes ReflectFact, a self-reflective agent framework that introduces a three-stage mechanism—explicit reasoning path planning, evidence drift verification, and reflective reasoning validation—to automatically detect and correct positional and substitution biases within reasoning chains for the first time. By integrating multi-hop question decomposition, evidence-grounded re-answering, step-wise consistency checking, and verification chain aggregation, ReflectFact effectively mitigates the misalignment between subtasks and the overarching verification goal. The method achieves state-of-the-art performance, surpassing the strongest baseline by 3.32% on HOVER and 2.78% on EX-FEVER.
为解决长篇叙事材料分析难题,提出ClueWeaver框架,通过奖励引导的双代理机制实现证据检索与解释,优化紧凑型本地模型性能。
本文提出MulVec方法,通过细粒度角色感知匹配解决无需训练的零样本组合图像检索问题,提高检索精度。
Existing multi-hop fact verification methods often suffer from reasoning deviations or erroneous conclusions due to a lack of global objective awareness and conflicts between parametric knowledge and retrieved evidence. To address these issues, this work proposes ReflectFact, a self-reflective agent framework that introduces a three-stage mechanism—explicit reasoning path planning, evidence drift verification, and reflective reasoning validation—to automatically detect and correct positional and substitution biases within reasoning chains for the first time. By integrating multi-hop question decomposition, evidence-grounded re-answering, step-wise consistency checking, and verification chain aggregation, ReflectFact effectively mitigates the misalignment between subtasks and the overarching verification goal. The method achieves state-of-the-art performance, surpassing the strongest baseline by 3.32% on HOVER and 2.78% on EX-FEVER.