EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use

πŸ“… 2026-02-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge in online video understanding posed by the tension between the unbounded visual stream and the limited context window of multimodal large language models, which hinders simultaneous modeling of long-range dependencies and fine-grained details. To this end, the authors propose EventMemAgent, an active agent framework grounded in hierarchical event-centric memory. It constructs short-term memory through event boundary detection and reservoir sampling, structures long-term memory into a searchable archive, and integrates a multi-granularity perception toolkit. Furthermore, agentic reinforcement learning is introduced to internalize tool invocation and reasoning as part of the agent’s policy. The method achieves competitive performance across multiple online video understanding benchmarks and represents the first end-to-end joint optimization of hierarchical event-granular memory and active perception.

Technology Category

Application Category

πŸ“ Abstract
Online video understanding requires models to perform continuous perception and long-range reasoning within potentially infinite visual streams. Its fundamental challenge lies in the conflict between the unbounded nature of streaming media input and the limited context window of Multimodal Large Language Models (MLLMs). Current methods primarily rely on passive processing, which often face a trade-off between maintaining long-range context and capturing the fine-grained details necessary for complex tasks. To address this, we introduce EventMemAgent, an active online video agent framework based on a hierarchical memory module. Our framework employs a dual-layer strategy for online videos: short-term memory detects event boundaries and utilizes event-granular reservoir sampling to process streaming video frames within a fixed-length buffer dynamically; long-term memory structuredly archives past observations on an event-by-event basis. Furthermore, we integrate a multi-granular perception toolkit for active, iterative evidence capture and employ Agentic Reinforcement Learning (Agentic RL) to end-to-end internalize reasoning and tool-use strategies into the agent's intrinsic capabilities. Experiments show that EventMemAgent achieves competitive results on online video benchmarks. The code will be released here: https://github.com/lingcco/EventMemAgent.
Problem

Research questions and friction points this paper is trying to address.

online video understanding
streaming media
context window limitation
long-range reasoning
fine-grained details
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical memory
event-centric representation
online video understanding
agentic reinforcement learning
adaptive tool use
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
S
Siwei Wen
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
Z
Zhangcheng Wang
4Paradigm Inc.
X
Xingjian Zhang
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University
Lei Huang
Lei Huang
Postdoctoral Fellow, MGH/Harvard Medical School/Broad Institute/MIT (Email: layne_huang@outlook.com)
Machine LearningAI for ScienceDrug Discovery
W
Wenjun Wu
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University; Hangzhou International Innovation Institute, Beihang University