Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning
This work addresses the challenge of semantic inconsistency in existing gaze target estimation methods, which typically rely on pixel-level regression and struggle in complex scenes. The study reframes the task as an object-centric hierarchical reasoning problem and introduces a two-stage framework: the first stage identifies potential gazed objects, while the second stage achieves precise localization by integrating object semantics–guided feature alignment, multi-scale fusion, and geometric constraints derived from head pose and gaze direction. By explicitly modeling semantic entities and incorporating geometric priors, the proposed method achieves state-of-the-art performance with only 7.1M parameters, attaining AUC scores of 0.961, 0.948, 0.987, and 0.977 on the GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real datasets, respectively, significantly improving both semantic consistency and localization accuracy.