Tool Retrievers Are Underestimated: Annotation Expansion Reveals True Capability

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出ToolEX框架解决工具检索基准中一对一标注导致的评估偏差问题,通过自动发现和标注功能等价的工具组合,更准确地评估检索器性能。
📝 Abstract
In open-world scenarios with massive and evolving tool repositories, tool-augmented large language models rely on a retriever to surface relevant tools for a given query. Because such repositories often contain many tools that implement the same functionality, a single query can often be resolved by several distinct but functionally equivalent tool combinations, making the natural query-to-tool mapping inherently one-to-many. However, existing tool retrieval benchmarks annotate each query with a single relevant tool combination, collapsing this one-to-many mapping into a rigid one-to-one annotation and causing valid retrieved tools to be misjudged as failures. To address this, we propose ToolEX (Tool Equivalent eXpansion), a framework that automatically discovers and annotates the tool combinations functionally equivalent to the labeled ones. Applied to the 7,360-query Tool-DE benchmark, ToolEX finds that 67.9% of sub-queries admit equivalent alternatives, expanding the singular ground truth to an average of 5.3 valid combinations per query. Using the expanded benchmark ToolEQ, we re-evaluate eight base retrievers and two fine-tuned variants; metrics on ToolEQ rise substantially over Tool-DE, showing that one-to-one annotation systematically underestimates retrievers and that 30--47% of the reported fine-tuning gain is an evaluation artifact rather than genuine improvement. Applying the same pipeline to skill retrieval on SkillRet further confirms that the one-to-one problem extends beyond tool retrieval.
Problem

Research questions and friction points this paper is trying to address.

tool retrieval
one-to-many mapping
annotation
functional equivalence
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

ToolEX
equivalent tool combinations
annotation expansion
one-to-many mapping
retriever evaluation
🔎 Similar Papers
Y
Yanyu Zhu
Shenzhen International Graduate School, Tsinghua University
C
Chenheng Zhang
Peking University
S
Shaoshen Chen
Shenzhen International Graduate School, Tsinghua University
H
Hoilam Pao
Shenzhen International Graduate School, Tsinghua University
Y
Yufei Zhang
Meituan, Beijing
Jiajun Chai
Jiajun Chai
Meituan Inc.
Reinforcement LearningLLMsAgentic Learning
D
Dongnian Wang
Meituan, Beijing
Z
Zhaoyu Hu
Meituan, Beijing
Guojun Yin
Guojun Yin
Meituan, University of Science and Technology of China
MultimodalityComputer VisionFoundation ModelsDeep LearningImage/Video Processing
W
Wei Lin
Meituan, Beijing
H
Hai-Tao Zheng
Shenzhen International Graduate School, Tsinghua University