GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
Existing 4D scene graphs struggle to model the structured temporal dynamics of human activities, and current activity representations either decouple from persistent 3D scenes or rely on predefined event boundaries and object associations. This work proposes a spatiotemporal memory system that integrates a persistent 4D scene graph with a dual-layer human–object interaction framework and a goal-driven hierarchical event structure to automatically extract, associate, and cluster atomic interactions from RGB-D streams into structured events. Through hierarchical event modeling and a context-aware object association refinement mechanism, the system enables retrospective and inferable activity memory. It achieves performance of 0.71–0.75 on standard benchmarks and scores 0.73 and 0.75 on newly introduced Space2Event and Event2Space query tasks, respectively, approaching the upper bound attainable with ground-truth events and associations.