LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the visual evidence loss in video agents caused by text-only observation interfaces during long video understandingโ€”a limitation referred to as the tool observation bottleneck. To overcome this, the authors propose LAVE, a training-free framework that introduces a novel latent visual evidence reuse mechanism. LAVE employs a dual-channel observation interface to preserve visual information from completed tool invocations, incorporating a latent channel enriched with tool role, timestamp, and spatial coordinates. Coupled with an entropy-constrained temporal alignment routing strategy and retrieval-based evidence fusion, this approach enhances multi-step planning capabilities. Experiments demonstrate significant performance gains on Video-MME, LongVideoBench, and CG-Bench, with LAVE achieving a 3.76-point improvement in overall Video-MME score under the same frame budget.
๐Ÿ“ Abstract
Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this challenge by iteratively invoking visual Tools at different temporal scales, but their Tool-Planner communication typically relies on textual observations. Such text-only interfaces provide lossy summaries of Tool computations, causing previously computed visual evidence not verbalized to be discarded and unavailable for subsequent planning. We identify this limitation as the Tool observation bottleneck and propose Latent Visual Evidence-Enhanced Planning (LAVE), a training-free framework for reusing latent visual evidence from completed Tool calls. LAVE introduces a dual-channel observation interface: the visible channel preserves the original textual trajectory, while the latent channel stores pre-verbal visual updates with their Tool roles, source-frame timestamps, and visual locations. During planning, LAVE retrieves evidence relevant to the current Planner state but not covered by textual observations, and integrates it through bounded timestamp-aligned latent updates with entropy-constrained frame-time routing. This enables video agents to reuse existing visual computation without additional training, frame replay, or modifications to the original orchestration. Extensive experiments on Video-MME, LongVideoBench, and CG-Bench show that LAVE consistently improves video tool-use agents across backbones. Under a comparable frame budget, LAVE improves the Video-MME overall score by 3.76 points over the strongest baseline, demonstrating the effectiveness of latent visual evidence reuse for multi-step video-agent planning.
Problem

Research questions and friction points this paper is trying to address.

video tool-use agents
long-video understanding
visual evidence reuse
tool observation bottleneck
latent visual information
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent visual evidence
tool-use agents
dual-channel observation
timestamp-aligned routing
training-free framework
๐Ÿ”Ž Similar Papers
Zijian Wang
Zijian Wang
China University of Petroleum(East China)
RLLLMNLP
Junnan Zhu
Junnan Zhu
Institute of Automation Chinese Academy of Sciences
Natural Language Processing
R
Rongzhen Li
Chongqing National Data AI Research Institute, AI Research Lab
X
Xiao Liu
College of Computer Science, Chongqing University, Chongqing, China
G
Guohui Xiang
Chongqing National Data AI Research Institute, AI Research Lab
Q
Quan Lu
Chongqing National Data AI Research Institute, AI Research Lab
L
Lijia Liu
College of Computer Science, Chongqing University, Chongqing, China
Yining Wang
Yining Wang
NLP Reseacher, Unisound
Natural Language ProcessingMachine Translation
J
Jiang Zhong
College of Computer Science, Chongqing University, Chongqing, China
K
Kaiwen Wei
College of Computer Science, Chongqing University, Chongqing, China