🤖 AI Summary
Existing video understanding benchmarks fail to evaluate models’ long-term spatiotemporal memory over timescales spanning days to weeks. To address this gap, this work introduces the first month-scale egocentric video benchmark for long-term memory assessment, comprising over 300 hours of daily recordings from 20 participants and 1,443 carefully curated multiple-choice questions. The authors propose a 14-task evaluation framework grounded in three layers of cognitive capabilities: schematic integration, episodic indexing, and cascaded reasoning. Systematic evaluation reveals that even the strongest current multimodal large language model, Gemini 2.5 Pro, achieves only a macro-averaged accuracy of 71.8%, substantially below the human baseline of 94.2%, and performs near chance level on tasks such as path reasoning—highlighting its lack of genuine long-term memory and suggesting it functions more as a lossy summarizer than a true memory system.
📝 Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.