TransGaze-Object: Transformer Based Driver Gaze Object Prediction Framework in Real Driving

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出TransGaze-Object框架,通过Transformer模型结合驾驶员面部和交通场景信息直接预测注视对象,提高预测准确性和减少错误率。
📝 Abstract
Driver gaze provides information regarding driver visual attention and situational awareness to the surrounding traffic. Existing driver gaze estimation studies represent gaze in terms of gaze zone or gaze vector/point-of-gaze (PoG). However, object-level gaze information provides a more semantically meaningful representation of visual attention by identifying attended objects, such as vehicles, pedestrians, or traffic signals. In this study, we propose an end-to-end driver gaze object prediction framework, TransGaze-Object, Transformer-based Gaze Object prediction model. The proposed framework first extracts facial features, including face and iris-weighted eye features, along with trafficobject spatial features. A transformer based cross-attention mechanism is then used to compute similarity scores and attention weights for predicting the drivers gaze object. To train this model, we propose a benchmark driver gaze dataset, Urban Driving-Face Scene Gaze (UD-FSG), comprising synchronized driver-face and traffic-scene images, scene objects bounding boxes, and gaze labels in terms of 2D gaze coordinate and gaze object. The TransGaze-Object model achieves an overall accuracy of 60% for gaze-object prediction, compared to 51% accuracy obtained from associating the estimated Point-of-Gaze to traffic objects. The error analysis reveals that TransGaze-Object reduces confusion between traffic objects (predicted) and the background (ground-truth), achieving an error rate of 11.68%, a 49.7% relative reduction compared with 23.21% error obtained from PoG-based gaze-object association. Overall, the results demonstrate the effectiveness of directly predicting gaze objects from driver-face and traffic-scene information, rather than estimating an intermediate Point-of-Gaze and subsequently associating it with traffic objects.
Problem

Research questions and friction points this paper is trying to address.

driver gaze
object-level gaze information
visual attention
traffic objects
gaze object prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer
Driver Gaze Object Prediction
Cross-Attention Mechanism
UD-FSG Dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pavan Kumar Sharma
Department of Civil Engineering, Indian Institute of Technology Kanpur, Kanpur-208016, U.P., India
A
Ayush Pande
Department of Computer Science and Engineering, Indian Institute of Technology Kanpur, Kanpur-208016, U.P., India
P
Pranamesh Chakraborty
Department of Civil Engineering, Indian Institute of Technology Kanpur, Kanpur-208016, U.P., India