SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding
Existing approaches typically model human actions and gaze in isolation, overlooking their intrinsic coupling in behavior understanding. This work proposes SAGE, a unified framework that, for the first time, jointly models the recognition and future prediction of human-object interaction (HOI) and gaze within an end-to-end trainable architecture applicable to both egocentric and exocentric settings. Built upon Transformers, SAGE integrates gaze cues into spatiotemporal attention mechanisms to enable joint reasoning over current and future action-gaze dynamics. We also introduce Exo-Cook, the first dataset with synchronized HOI and gaze annotations for exocentric videos. Experiments demonstrate that SAGE achieves performance on par with or superior to state-of-the-art methods specialized for individual tasks across VidHOI, EGTEA Gaze+, and Exo-Cook benchmarks.