GazeDiT: Gaze-Accurate Diffusion Image Generation for Eye Tracking via Spatial Conditioning

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决眼动追踪中精确标签控制难题,提出GazeDiT模型,通过内部构建的空间条件生成与4D视线对应的图像,提高眼动追踪精度。
📝 Abstract
Diffusion models are increasingly used to generate synthetic training data, but precise label control remains difficult when the conditioning signal is low-dimensional and coarse. Text-conditioned images are judged by broad prompt consistency, whereas supervised training requires precise correspondence between each image and its numerical label. This is challenging in eye tracking, where a 4D binocular gaze is expressed through subtle, spatially localized pupil and iris geometry. We introduce GazeDiT, a diffusion model that generates images for a requested 4D gaze through an internally constructed spatial condition that grounds the global gaze label in this local geometry. During training, a frozen SegFormer extracts pupil/iris geometry from diverse real images, allowing the model to learn realistic appearance conditioned on that geometry. At inference, a physical eye renderer samples gaze-consistent geometries by varying anatomy and camera state, enabling diverse synthesis without a source image. GazeDiT achieves substantially lower tail gaze-label error than other diffusion baselines, approaching the error of the same frozen gaze estimator on real images. Its generated data also improves the downstream eye tracker, reducing gaze error on difficult cases from 3.05° to 2.80° in the smallest cohort.
Problem

Research questions and friction points this paper is trying to address.

eye tracking
4D gaze
synthetic data generation
diffusion models
label control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaze-Accurate
Spatial Conditioning
Diffusion Model
Eye Tracking
Synthetic Data Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.