🤖 AI Summary
Progress in consumer-grade electroencephalography–eye-tracking (EEG-ET) research has been hindered by the scarcity of high-quality, time-synchronized multimodal data. Method: This work introduces and publicly releases the first large-scale, multi-paradigm, standardized benchmark dataset for consumer-grade EEG-ET synchronization. It comprises 113 participants, 116 sessions, and 11.75 hours of high-fidelity synchronized recordings—acquired using low-cost commercial EEG systems (e.g., OpenBCI) and webcam-based eye trackers—spanning four oculomotor paradigms: saccades, smooth pursuit, fixation, and free viewing. All data undergo rigorous temporal alignment, bandpass filtering, and missing-value imputation; accompanying open-source preprocessing and analysis code is provided. Contribution/Results: The dataset substantially lowers hardware barriers for EEG-ET research, enhances reproducibility, and provides critical empirical support for gaze decoding under challenging conditions—including low-light environments and camera-free settings.
📝 Abstract
Electroencephalography-based eye tracking (EEG-ET) leverages eye movement artifacts in EEG signals as an alternative to camera-based tracking. While EEG-ET offers advantages such as robustness in low-light conditions and better integration with brain-computer interfaces, its development lags behind traditional methods, particularly in consumer-grade settings. To support research in this area, we present a dataset comprising simultaneous EEG and eye-tracking recordings from 113 participants across 116 sessions, amounting to 11 hours and 45 minutes of recordings. Data was collected using a consumer-grade EEG headset and webcam-based eye tracking, capturing eye movements under four experimental paradigms with varying complexity. The dataset enables the evaluation of EEG-ET methods across different gaze conditions and serves as a benchmark for assessing feasibility with affordable hardware. Data preprocessing includes handling of missing values and filtering to enhance usability. In addition to the dataset, code for data preprocessing and analysis is available to support reproducibility and further research.