VLX-VR: An Agentic-Aware Video Reasoning Model

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出VLX-VR模型,通过强化学习训练以解决视频理解中证据获取和利用的问题,实现对多模态数据的高效推理。
📝 Abstract
Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete, ambiguous, or conflicting. We present VLX-VR, an agentic-aware video reasoning model trained within a video reasoning framework defined by a Think--Memory--Observation loop. At each step, VLX-VR determines the needed evidence, invokes read_memory or write_memory, incorporates the returned Observation, and decides whether to continue or produce the task output. We train VLX-VR with multimodal data, including videos and agent trajectories, using reinforcement learning to learn evidence acquisition, memory use, and termination. On MINERVA, VLX-VR achieves state-of-the-art performance among the models included in our comparison, with 78.79% accuracy. Under the original three duration groups, its accuracies are 76.70%, 78.73%, and 80.92%, with a cross-duration accuracy variance of 2.97~$\mathrm{pp}^2$. On correctly answered samples, 96.20% of VLX-VR's reasoning traces are consistent with the MINERVA reference reasoning traces and the evidence described by them, while approximately 75.80% of all evaluated samples satisfy both answer correctness and this evidence-grounded trace criterion. These results show strong performance and broadly stable behavior across durations, while counting, state changes, causal reasoning, and spatial perception remain challenging.
Problem

Research questions and friction points this paper is trying to address.

video understanding
evidence acquisition
adaptive reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic-aware
video reasoning
reinforcement learning
multimodal data
memory use
S
Sheng Li
Om AI Research
P
Peng Liu
Om AI Research
Q
Qianqian Zhang
Om AI Research
T
Tiancheng Zhao
Om AI Research