AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AdaVDR,通过自适应工具调用和反思解决视频深度研究中因工具使用不当导致的问题,利用数据构建管道生成高质量QA对,并通过强化学习优化模型性能。
📝 Abstract
Video deep research answers complex questions by jointly understanding video content and retrieving external knowledge from the open Web. However, diverse questions and videos require different tool-use strategies, and inappropriate tool calls can produce incorrect results. Uncertain grounding and retrieval also make unnecessary interactions costly and error-prone, increasing latency and reasoning errors. To address these challenges, we propose AdaVDR, an adaptive video deep research agent with adaptive tool invocation and reflection. AdaVDR selects tools according to the task and its capabilities, and backtracks only when unreliable intermediate results require correction. To enable these capabilities, we develop a video deep research data construction pipeline. We first discover retrieval-relevant events and entities in diverse videos and acquire detailed information through grounding and external retrieval to construct high-quality QA pairs. For each QA, task-specific prompts organize the information acquisition process into a tool-use trajectory, allowing different question and video types to follow different grounding and retrieval strategies. We further introduce model-conditioned tool necessity filtering, which evaluates tool calls against the target model's video understanding and internal knowledge, removing tools or tool chains the model can bypass. This yields trajectories tailored to the target model's video understanding capability and knowledge. Using this pipeline, we construct training data and VDR-EE, a benchmark covering entity-centric and event-centric questions. We perform supervised fine-tuning followed by reinforcement learning with a redundancy-aware reward to strengthen adaptive tool invocation and reflection. Experiments show that our method performs best among the evaluated open-source models on VDR-EE and substantially improves over its base models on VideoDR.
Problem

Research questions and friction points this paper is trying to address.

Video Deep Research
Tool Use Strategies
Retrieval Errors
Reasoning Errors
Adaptive Tool Invocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive tool invocation
reflection
video deep research data construction pipeline
model-conditioned tool necessity filtering
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
X
Xintong Zhang
Accio Team, Alibaba Group
Xiaomeng Fan
Xiaomeng Fan
Beijing Institute of Technology
machine learningcomputer vision
Shilin Yan
Shilin Yan
Fudan University
MLLMsComputer VisionMulti-Modal
E
Ekko He
Accio Team, Alibaba Group
Z
Zicheng Liu
Accio Team, Alibaba Group
Z
Zijian Zou
Accio Team, Alibaba Group
G
Guannan Zhang
Accio Team, Alibaba Group
Yuwei Wu
Yuwei Wu
Ph.D. candidate, GRASP Lab, University of Pennsylvania
RoboticsTrajectory OptimizationTask and Motion Planning
Z
Zhi Gao
Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology
Hongwei Xue
Hongwei Xue
University of Science and Technology of China
Multi-ModalVision-Language