Eyes on Target: Gaze-Aware Object Detection in Egocentric Video

📅 2025-11-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the neglect of human visual attention in first-person video object detection, this paper proposes a gaze-guided Vision Transformer framework. It models eye-tracking trajectories as dynamic attention priors and injects them into the multi-head self-attention mechanism; a gaze-aware attention head importance metric is further introduced to explicitly modulate each head’s response strength to fixated regions. Additionally, depth information is fused to enhance spatial perception. The method achieves significant improvements over non-gaze-aware baselines—up to +3.2–5.7% mAP—on a custom simulator dataset and public benchmarks including Ego4D Ego-Motion and Ego-CH-Gaze. Ablation studies validate the effectiveness of both the gaze-guidance mechanism and the depth-fusion module. This work is the first to systematically demonstrate the interpretability value of eye movement signals for evaluating first-person object detection, establishing a novel paradigm for attention modeling in embodied intelligence.

Technology Category

Application Category

📝 Abstract
Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric videos. Our approach injects gaze-derived features into the attention mechanism of a Vision Transformer (ViT), effectively biasing spatial feature selection toward human-attended regions. Unlike traditional object detectors that treat all regions equally, our method emphasises viewer-prioritised areas to enhance object detection. We validate our method on an egocentric simulator dataset where human visual attention is critical for task assessment, illustrating its potential in evaluating human performance in simulation scenarios. We evaluate the effectiveness of our gaze-integrated model through extensive experiments and ablation studies, demonstrating consistent gains in detection accuracy over gaze-agnostic baselines on both the custom simulator dataset and public benchmarks, including Ego4D Ego-Motion and Ego-CH-Gaze datasets. To interpret model behaviour, we also introduce a gaze-aware attention head importance metric, revealing how gaze cues modulate transformer attention dynamics.
Problem

Research questions and friction points this paper is trying to address.

Enhancing object detection using human gaze in egocentric videos
Injecting gaze features into Vision Transformer attention mechanisms
Improving detection accuracy in task-critical simulation scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaze-guided object detection in egocentric videos
Injecting gaze features into Vision Transformer attention
Depth-aware framework biasing feature selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Vishakka Lall
Singapore Polytechnic, Centre of Excellence in Maritime Safety, Singapore
Y
Yisi Liu
Singapore Polytechnic, Centre of Excellence in Maritime Safety, Singapore