TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出TRACE方法,通过构建证据包来解决长视频理解中答案基于不完全观察的问题,提高了视频理解和答案准确性。
📝 Abstract
A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness or predicted evidence intervals, but the frames a method decodes before answering are rarely audited, so correct answers can still rest on incomplete observation. We introduce VES-Bench, a 600-question benchmark of Temporal Ordering and Event Counting items over 348 public long videos. Each item carries a jointly necessary set of evidence intervals, letting us audit at three strictness levels whether a method's decoded frames cover every one of them. We also propose TRACE, a training-free agent that grounds answers in raw visual clips, builds an evidence bundle round by round, and stops only when the answer stabilises as the bundle grows and a final pass over the same clips returns the same answer. Under a same-backbone audit, TRACE answers 50.7% of questions correctly with at least two decoded frames inside every evidence interval, at 98.7 frames per question: over 10 points above uniform decoding at 128 frames (40.2%), and within 2.6 points of uniform decoding at 256 frames at 0.39x its frame cost, while reaching the highest answer accuracy in the audit (63.5%). TRACE also stays competitive on Video-MME (86.1), LVBench (75.6), and LongVideoBench (75.1).
Problem

Research questions and friction points this paper is trying to address.

long-video understanding
evidence intervals
frame decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal Retrieval
Anchored and Convergent Evidence
Long-Horizon Video Understanding
VES-Bench
TRACE
💼 Related Jobs
No related jobs found.
P
Pengyiang Liu
Colab, Beihang University
Junbo Niu
Junbo Niu
Peking University
Foundation Model
X
Xiaoyang Hu
Colab, Beihang University
Z
Zhongyue Shi
Colab, Beihang University
Z
Zitian Wang
Colab, Beihang University
Linjiang Huang
Linjiang Huang
BUAA<<CUHK<<CASIA
Computer VisionPattern RecognitionMachine Learning
Si Liu
Si Liu
Fred Hutchinson Cancer Center
GenomicsBiostatisticsAnomaly DetectionOpen Category Detection