Institution profile

Qingdao University of Science and Technology

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Camera-based implicit mind reading by capturing higher-order semantic dynamics of human gaze within environmental context

Jul 17, 2025

Existing emotion recognition methods suffer from key limitations: overt behavioral cues (e.g., facial expressions, speech) are easily feigned; physiological signals require invasive instrumentation; and gaze analysis often neglects environmental context. To address these, we propose a non-intrusive, continuous emotion recognition paradigm leveraging only a standard high-definition camera to simultaneously capture naturalistic gaze trajectories and head motion. For the first time, our approach deeply integrates gaze dynamics, environmental semantics, and temporal evolution into a unified spatial–semantic–temporal behavioral model. It operates implicitly—requiring no user cooperation or specialized sensors—to decode affective states. Experimental results demonstrate high robustness, real-time performance, low cost, and strong scalability in unconstrained real-world settings. This work advances computational affective science by formalizing emotion as an emergent product of human–environment interaction, offering a novel, scalable, and ecologically valid framework for implicit affective computing.

0 citationsRead paper

TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection

Mar 18, 2025

In video object detection, unimodal models—whether CNN- or ViT-based—struggle to jointly capture fine-grained local details and holistic spatiotemporal context. To address this, we propose a dual-path collaborative architecture: one path employs a spatiotemporal Transformer to model long-range temporal dependencies; the other adopts a graph-enhanced GraphFormer that explicitly encodes intra-frame object relationships and inter-frame correspondences. We further introduce a novel adaptive global-local feature fusion module that dynamically couples the complementary strengths of these two paradigms. The entire framework operates end-to-end on video frame sequences. Evaluated on ImageNet VID, our method achieves 86.5% mAP while maintaining real-time inference at 41.0 FPS on a single A100 GPU—setting a new state-of-the-art performance.

0 citationsRead paper
Recent publications

Latest Papers

Camera-based implicit mind reading by capturing higher-order semantic dynamics of human gaze within environmental context

Jul 17, 2025

Existing emotion recognition methods suffer from key limitations: overt behavioral cues (e.g., facial expressions, speech) are easily feigned; physiological signals require invasive instrumentation; and gaze analysis often neglects environmental context. To address these, we propose a non-intrusive, continuous emotion recognition paradigm leveraging only a standard high-definition camera to simultaneously capture naturalistic gaze trajectories and head motion. For the first time, our approach deeply integrates gaze dynamics, environmental semantics, and temporal evolution into a unified spatial–semantic–temporal behavioral model. It operates implicitly—requiring no user cooperation or specialized sensors—to decode affective states. Experimental results demonstrate high robustness, real-time performance, low cost, and strong scalability in unconstrained real-world settings. This work advances computational affective science by formalizing emotion as an emergent product of human–environment interaction, offering a novel, scalable, and ecologically valid framework for implicit affective computing.

0 citationsRead paper

TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection

Mar 18, 2025

In video object detection, unimodal models—whether CNN- or ViT-based—struggle to jointly capture fine-grained local details and holistic spatiotemporal context. To address this, we propose a dual-path collaborative architecture: one path employs a spatiotemporal Transformer to model long-range temporal dependencies; the other adopts a graph-enhanced GraphFormer that explicitly encodes intra-frame object relationships and inter-frame correspondences. We further introduce a novel adaptive global-local feature fusion module that dynamically couples the complementary strengths of these two paradigms. The entire framework operates end-to-end on video frame sequences. Evaluated on ImageNet VID, our method achieves 86.5% mAP while maintaining real-time inference at 41.0 FPS on a single A100 GPU—setting a new state-of-the-art performance.

0 citationsRead paper