Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出基于运动的标记化方法解决跨数据集的第一人称视线建模问题,通过比较不同表示方法评估其在目标预测和迁移上的表现。
📝 Abstract
Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse event labels are easy to model but can discard local motion structure. We formulate event-aligned, fixed-horizon angular displacement as an interpretable, event-conditioned motion vocabulary and compare it with event-only, spatial, absolute-angle, learned vector-quantized, and continuous representations. To assess transfer alongside target predictability and token collapse, our evaluation combines next-token prediction with target-domain regret, low-order target references, paired bootstrap, order sensitivity, motif overlap, and frozen structural probes. In an event-aligned headset benchmark, angular-motion tokens have lower target-domain regret than frozen-codebook VQ tokens in one transfer direction, while the reverse direction is inconclusive. The probes reveal complementary representation properties, and event-only tokens show that low perplexity can retain little motion information. On a third egocentric dataset, a matched comparison of I-VT, native, and frame-span interfaces shows that event construction materially changes transfer: native events have the lowest regret into EGTEA, while frame-span events have zero motif overlap and fail severely as a source. Motion-based tokenization therefore provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.
Problem

Research questions and friction points this paper is trying to address.

Cross-Dataset
Egocentric Gaze Modeling
Tokenization
Motion Representation
Event-Conditioned
Innovation

Methods, ideas, or system contributions that make the work stand out.

event-aligned angular displacement
motion-based tokenization
cross-dataset transfer
gaze modeling
💼 Related Jobs
No related jobs found.
V
Virmarie Maquiling
Human-Centered Technologies for Learning, Technical University of Munich (TUM), Munich, Germany; Munich Center for Machine Learning (MCML), Munich, Germany
Zhuojiang Cai
Zhuojiang Cai
Technical University of Munich
Human-Computer InteractionComputer Vision
Enkelejda Kasneci
Enkelejda Kasneci
Professor at the Technical University of Munich
Eye TrackingAI in EducationHuman-Centered AIComputational InteractionHCI