Institution profile

Wakayama University

Academic institutionasia · jp
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning Reactive Human Motion Generation from Paired Interaction Data Using Transformer-Based Models

Apr 23, 2026

This study addresses the challenge of generating reactive and structurally consistent human motion sequences for one participant based on the actions of another in dyadic interaction scenarios. To this end, the authors construct a paired action–reaction motion dataset derived from boxing match videos and propose a Transformer-based architecture augmented with character ID embeddings to explicitly distinguish individual identities, thereby enhancing interaction awareness and structural consistency. The authors systematically evaluate several Transformer variants—including standard Transformer, iTransformer, and Crossformer—and find that the standard Transformer demonstrates superior stability in long-term generation, effectively avoiding pose collapse. Moreover, the incorporation of character ID embeddings significantly mitigates structural degradation and improves motion coherence. This work represents the first effort to integrate character ID embeddings into dyadic motion generation, offering a novel approach to interactive human motion synthesis.

0 citationsRead paper

DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners

Dec 22, 2025

Current objective assessment of speech intelligibility for hearing-impaired listeners relies heavily on clean reference signals—a major bottleneck in clinical and hearing-aid fitting scenarios. To address this, we propose DeepGESI, the first fully non-intrusive deep learning model for reference-free prediction of the hearing-loss-specific metric GESI (Generalized Estimation of Speech Intelligibility). DeepGESI takes only distorted speech as input and performs end-to-end regression, jointly modeling time-frequency acoustic representations and hearing-loss perception priors. Unlike conventional reference-dependent methods, DeepGESI enables pure no-reference GESI estimation, significantly enhancing practicality in real-world applications. Evaluated on the CPC2 dataset, it achieves high correlation with human-rated GESI (Spearman ρ > 0.92) and accelerates inference by over 20× compared to prior approaches. This work establishes a new paradigm for objective, efficient, and personalized speech intelligibility assessment tailored to hearing impairment.

0 citationsRead paper
Recent publications

Latest Papers

Learning Reactive Human Motion Generation from Paired Interaction Data Using Transformer-Based Models

Apr 23, 2026

This study addresses the challenge of generating reactive and structurally consistent human motion sequences for one participant based on the actions of another in dyadic interaction scenarios. To this end, the authors construct a paired action–reaction motion dataset derived from boxing match videos and propose a Transformer-based architecture augmented with character ID embeddings to explicitly distinguish individual identities, thereby enhancing interaction awareness and structural consistency. The authors systematically evaluate several Transformer variants—including standard Transformer, iTransformer, and Crossformer—and find that the standard Transformer demonstrates superior stability in long-term generation, effectively avoiding pose collapse. Moreover, the incorporation of character ID embeddings significantly mitigates structural degradation and improves motion coherence. This work represents the first effort to integrate character ID embeddings into dyadic motion generation, offering a novel approach to interactive human motion synthesis.

0 citationsRead paper

DeepGESI: A Non-Intrusive Objective Evaluation Model for Predicting Speech Intelligibility in Hearing-Impaired Listeners

Dec 22, 2025

Current objective assessment of speech intelligibility for hearing-impaired listeners relies heavily on clean reference signals—a major bottleneck in clinical and hearing-aid fitting scenarios. To address this, we propose DeepGESI, the first fully non-intrusive deep learning model for reference-free prediction of the hearing-loss-specific metric GESI (Generalized Estimation of Speech Intelligibility). DeepGESI takes only distorted speech as input and performs end-to-end regression, jointly modeling time-frequency acoustic representations and hearing-loss perception priors. Unlike conventional reference-dependent methods, DeepGESI enables pure no-reference GESI estimation, significantly enhancing practicality in real-world applications. Evaluated on the CPC2 dataset, it achieves high correlation with human-rated GESI (Spearman ρ > 0.92) and accelerates inference by over 20× compared to prior approaches. This work establishes a new paradigm for objective, efficient, and personalized speech intelligibility assessment tailored to hearing impairment.

0 citationsRead paper