EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low token efficiency in first-person video understanding caused by the absence of eye-tracking hardware in consumer-grade smart glasses. To overcome this limitation, we propose EgoGazeLite, a lightweight dual-stream gaze predictor that substitutes dedicated sensors with on-device real-time gaze estimation and visual token pruning to enable efficient multimodal large model inference. Experimental results demonstrate that the proposed model, comprising only 15.7M parameters, achieves real-time inference at 21.6 ms per frame. Notably, its hardware-free gaze-guided cropping yields performance statistically indistinguishable from ground-truth annotations. Consequently, EgoGazeLite effectively satisfies stringent edge deployment constraints while significantly enhancing video understanding efficacy, offering a viable software-centric solution for resource-constrained wearable computing.
📝 Abstract
The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget. Memory and compute cost scale with the number of visual tokens, and high-resolution video quickly becomes expensive to transmit and process at scale. Prior work (GazeLLM) addresses this by cropping the video around the camera wearer's gaze. This reduces the number of visual tokens by about tenfold while maintaining or improving the quality of full-resolution descriptions. However, this compression strategy depends on dedicated eye-tracking hardware, which is unavailable on consumer smart glasses. Building a software-only substitute poses a joint constraint: the predictor must be accurate enough to preserve downstream description quality, yet light enough to run on-device, within the power and compute budget of a smartphone. We address this with EgoGazeLite, a lightweight dual-process gaze predictor for egocentric video. Across two MLLMs, three automated metrics, and two LLM judges, predicted-gaze crops show no significant difference from ground-truth-gaze crops. Equivalence is confirmed in all ten cases. EgoGazeLite achieves this at 15.7M parameters, 6.71 GFLOPs, and runs the full gaze-and-crop pipeline end-to-end in real time (21.6 ms/frame) on consumer accelerator hardware. Together, these results remove the need for eye-tracking hardware for token-efficient, gaze-conditioned egocentric video understanding with MLLMs.
Problem

Research questions and friction points this paper is trying to address.

Egocentric video understanding
Multimodal LLMs
Token efficiency
On-device gaze prediction
Wearable devices
Innovation

Methods, ideas, or system contributions that make the work stand out.

Egocentric Gaze Prediction
Token-Efficient MLLM
Lightweight On-Device Model
Software-Only Eye Tracking
Real-Time Video Cropping
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Matteo Stoiber
Copenhagen Business School, Department of Digitalization, Copenhagen, Denmark
N
Niels Buus Lassen
Copenhagen Business School, Department of Digitalization, Copenhagen, Denmark