CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出CAFE框架,通过交替优化搜索代理和批评者角色,利用共同参数模型解决传统方法无法及时纠正错误的问题,提高搜索效率和准确性。
📝 Abstract
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles. CAFE initializes feedback-conditioned recovery from trajectories built around the base agent's own failures, then couples online and offline optimization. During online RL, a comparative feedback estimate uses a prompt-level call--skip success gap to shape request returns, while feedback-aware advantage shaping reweights token advantages before and after feedback. Offline, rollout-derived preference optimization learns feedback from matched successful and unsuccessful trajectories. On seven agentic search benchmarks, CAFE outperforms the evaluated RL-based search agents on average, retains its gains across all six out-of-domain benchmarks, and reduces answer-level hallucinations. One-sided ablations show that improving only the agent or only the critic eventually plateaus, whereas alternating the two updates continues to improve performance. These findings suggest that a self-improving search agent needs feedback that co-evolves with the policy it guides.
Problem

Research questions and friction points this paper is trying to address.

search agents
feedback
trajectory
error correction
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

CAFE
Co-Evolving Feedback
Search Agents
Online and Offline Optimization
Self-Improvement
Boyang Liu
Boyang Liu
Fudan University
NLPLLM reasoningEfficient AIReinforcement Learning
Senjie Jin
Senjie Jin
Fudan University
natural language processing
Peixin Wang
Peixin Wang
East China Normal University
Formal MethodsTrustworthy AIProgram Verification
Z
Zhangyue Yin
2LLM Department, Tencent
Y
Yibo Wang
2LLM Department, Tencent
Y
Yuhao Zhou
1Fudan University 2LLM Department, Tencent
X
Xinbing Liang
2LLM Department, Tencent
S
Shizheng Zhu
1Fudan University
Y
Yuhui Wang
1Fudan University
J
Jingqi Tong
1Fudan University
Zhiheng Xi
Zhiheng Xi
Fudan University
LLM ReasoningLLM-based Agents
Jiazheng Zhang
Jiazheng Zhang
Fudan University
Large Language ModelNatural Language ProcessingData Mining
C
Clive Bai
2LLM Department, Tencent
C
Clarenceai
2LLM Department, Tencent
B
Blaze Chen
2LLM Department, Tencent
T
Tao Gui
1Fudan University
Qi Zhang
Qi Zhang
Fudan University
SAGINsatellite routing
X
Xuanjing Huang
1Fudan University