DualStake: Dual-Path Confidence Calibration in Deep Research Agents

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决深度研究代理的过度自信问题,提出DualStake方法,通过双重路径校准E-Conf和A-Conf以提高信心表达准确性。
📝 Abstract
Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and decision-oriented generation. However, these agents suffer from severe overconfidence, making their expressed confidence unreliable for user trust and downstream abstention. To address this, we augment the Deep Research pipeline with step confidence elicitation after each retrieval, building on the commonly used post-answer verbalized confidence. Interestingly, we find that Evidence Confidence (E-Conf), elicited after the final retrieval step, provides a stronger uncertainty signal than Answer Confidence (A-Conf), elicited after answer generation, and that A-Conf is largely shaped by E-Conf. Based on these findings, we propose DualStake, a dual-path calibration method that applies margin-clipped, confidence-dependent stake rewards to jointly align E-Conf and A-Conf with answer correctness while limiting extreme confidence optimization. Experiments on Qwen2.5-7B, Qwen2.5-7B-Instruct, and Qwen3-4B across 8 QA benchmarks demonstrate that DualStake consistently improves calibration without sacrificing answer accuracy. The code is available at https://github.com/FloXXXt/DualStake.
Problem

Research questions and friction points this paper is trying to address.

Deep Research Agents
Overconfidence
Confidence Calibration
Multi-round Retrieval
Decision-oriented Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Path Calibration
Evidence Confidence (E-Conf)
Answer Confidence (A-Conf)
Margin-Clipped Stake Rewards
💼 Related Jobs
No related jobs found.
Y
Yinuo Xu
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences (CASIA)
Y
Yuwei Liang
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences (CASIA); School of Artificial Intelligence, University of Chinese Academy of Sciences
J
Jianjie Cheng
Meituan Inc.
M
Meng Wang
Meituan Inc.
Yongcan Yu
Yongcan Yu
Master Student, CASIA
Trustworthy AISafety in AI
S
Shuo Lu
NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences (CASIA)
Jian Liang
Jian Liang
Kuaishou Inc.
transfer learninggraph learning