Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出NTEP和NTEP-R方法,通过明确指定关键工具调用路径及奖励机制,解决视觉-语言模型在处理复杂查询时证据获取与利用不足的问题。
📝 Abstract
Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropriate tools are called, models often fail to extract the necessary information from the resulting observations. To address these limitations, we introduce the NTEP (Necessary Tool-Evidence Path), a novel annotation scheme that explicitly specifies the essential external evidence and corresponding tool calls for each query. Building upon this, we propose NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution. Specifically, our approach rewards the agent for aligning its pre-call intent with a necessary evidence-seeking goal, and for ensuring the information summarized from the post-call observation aligns with the necessary evidence. Furthermore, we introduce a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals. Extensive evaluations on seven image-grounded benchmarks demonstrate that our 8B-parameter instantiation, NTEP-8B, significantly improves both search-oriented accuracy and tool-use efficiency within a unified three-tool framework. These results highlight the critical value of fine-grained tool-evidence path supervision for training robust agentic VLMs.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
tool use
evidence acquisition
complex queries
external knowledge
Innovation

Methods, ideas, or system contributions that make the work stand out.

NTEP
NTEP-R
non-repeated-goal regularizer
Xingming Long
Xingming Long
中国科学院计算技术研究所
Y
Yu Liu
Institute of Information Engineering, Chinese Academy of Sciences
Z
Zhiwei Yang
Institute of Information Engineering, Chinese Academy of Sciences
H
Hanqi Feng
Department of Machine Learning, Carnegie Mellon University
S
Shaojie Zhang
MiLM Plus, Xiaomi Inc.
Barnabas Poczos
Barnabas Poczos
Associate professor, Carnegie Mellon University
Artificial IntelligenceMachine LearningStatistics
C
Chao Jiang
MiLM Plus, Xiaomi Inc.
Zhenbo Luo
Zhenbo Luo
XiaoMi
Vision Language ModelComputer Vision
Lei Jiang
Lei Jiang
Technical Institute of Physics and Chemistry, Chinese Academy of Sciences
bio-inspired interfacial materials with superwettability
P
Pei Fu
MiLM Plus, Xiaomi Inc.