From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
This paper addresses the scarcity of strong supervision in video salient object detection (VSOD) by proposing a novel weakly supervised approach leveraging eye-tracking signals. The method introduces three key components: (1) a Position-Semantic Embedding (PSE) module that jointly models spatial gaze distributions and semantic priors; (2) a Semantic-Local Query competition (SLQ) mechanism that dynamically enhances representations of salient regions; and (3) an Intra-Inter Mixed Contrastive (IIMC) learning paradigm that enforces consistency of weak supervision at both intra-frame and inter-frame granularities. Evaluated on five mainstream VSOD benchmarks, the framework consistently outperforms existing methods, achieving state-of-the-art performance across multiple metrics. Comprehensive experiments demonstrate its effectiveness, robustness, and generalization capability under weak supervision.