EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input
This study addresses the low token efficiency in first-person video understanding caused by the absence of eye-tracking hardware in consumer-grade smart glasses. To overcome this limitation, we propose EgoGazeLite, a lightweight dual-stream gaze predictor that substitutes dedicated sensors with on-device real-time gaze estimation and visual token pruning to enable efficient multimodal large model inference. Experimental results demonstrate that the proposed model, comprising only 15.7M parameters, achieves real-time inference at 21.6 ms per frame. Notably, its hardware-free gaze-guided cropping yields performance statistically indistinguishable from ground-truth annotations. Consequently, EgoGazeLite effectively satisfies stringent edge deployment constraints while significantly enhancing video understanding efficacy, offering a viable software-centric solution for resource-constrained wearable computing.