EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video understanding benchmarks fail to evaluate models’ long-term spatiotemporal memory over timescales spanning days to weeks. To address this gap, this work introduces the first month-scale egocentric video benchmark for long-term memory assessment, comprising over 300 hours of daily recordings from 20 participants and 1,443 carefully curated multiple-choice questions. The authors propose a 14-task evaluation framework grounded in three layers of cognitive capabilities: schematic integration, episodic indexing, and cascaded reasoning. Systematic evaluation reveals that even the strongest current multimodal large language model, Gemini 2.5 Pro, achieves only a macro-averaged accuracy of 71.8%, substantially below the human baseline of 94.2%, and performs near chance level on tasks such as path reasoning—highlighting its lack of genuine long-term memory and suggesting it functions more as a lossy summarizer than a true memory system.
📝 Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.
Problem

Research questions and friction points this paper is trying to address.

long-term spatiotemporal memory
egocentric video
video understanding benchmark
multimodal large language models
memory consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Egocentric Video
Long-Term Spatiotemporal Memory
Multimodal Large Language Models
Video Benchmark
Cognitive Evaluation Framework
🔎 Similar Papers
Weitao Chen
Weitao Chen
Unknown affiliation
H
Hu Jiaxin
Nanjing University, China
X
Xie Tianyidan
Nanjing University, China
Y
Yang Li
Nanjing University, China
Y
Yuyi Qian
Nanjing University, China
B
Banghao Xu
Nanjing University, China
Z
Ziheng Tang
Nanjing University, China
S
Shenyi Wang
Nanjing University, China
M
Mingyue Yu
Nanjing University, China
D
Duo Li
Nanjing University, China
J
Jiacheng Shi
Nanjing University, China
Gao Wang
Gao Wang
Assistant Professor at Columbia University Vagelos College of Physicians and Surgeons
Computational genomics
Zhan Xu
Zhan Xu
Unknown affiliation
computer graphicscomputer vision
Z
Zhicheng Qiu
Huawei Technologies Co., Ltd., China
X
Xuanfu Li
Huawei Technologies Co., Ltd., China
J
Jian Yang
Nanjing University, China
L
Lanjun Wang
Tianjin University, China
Z
Zili Yi
Nanjing University, China