Camera-based implicit mind reading by capturing higher-order semantic dynamics of human gaze within environmental context
Existing emotion recognition methods suffer from key limitations: overt behavioral cues (e.g., facial expressions, speech) are easily feigned; physiological signals require invasive instrumentation; and gaze analysis often neglects environmental context. To address these, we propose a non-intrusive, continuous emotion recognition paradigm leveraging only a standard high-definition camera to simultaneously capture naturalistic gaze trajectories and head motion. For the first time, our approach deeply integrates gaze dynamics, environmental semantics, and temporal evolution into a unified spatial–semantic–temporal behavioral model. It operates implicitly—requiring no user cooperation or specialized sensors—to decode affective states. Experimental results demonstrate high robustness, real-time performance, low cost, and strong scalability in unconstrained real-world settings. This work advances computational affective science by formalizing emotion as an emergent product of human–environment interaction, offering a novel, scalable, and ecologically valid framework for implicit affective computing.