Institution profile

Asahi Shimbun Company

Industry researchasia · jp
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

Jun 16, 2026

This work addresses the challenge of balancing model efficiency and performance in high-fidelity speech enhancement by proposing QC-GAN, a novel framework that integrates quaternion representations with the Conformer architecture for the first time. By leveraging Hamiltonian products to jointly model magnitude and phase in a structured manner, QC-GAN preserves their intrinsic correlation while substantially reducing parameter count. The approach further incorporates the MetricGAN training strategy and a metric learning-based discriminator to optimize perceptual quality. On the VoiceBank+DEMAND dataset, the model achieves a PESQ score of 3.48 with only 0.89 million parameters, and even a compact 35K-parameter variant attains 3.23—significantly outperforming conventional methods. Strong generalization capability is also demonstrated on the DNS-Challenge 3 benchmark.

0 citationsRead paper

Quaternion Self-Attention with Shared Scores

May 24, 2026

Existing quaternion self-attention mechanisms compute attention scores independently for each quaternion component, resulting in high computational overhead and inconsistent attention distributions. This work proposes a shared-score quaternion self-attention mechanism that generates a single real-valued attention score via quaternion inner product and shares the resulting attention distribution across all components. Theoretical analysis reveals that when queries and keys are pre-mixed through quaternion linear projections, component-wise independent scoring and shared scoring operate within the same interaction subspace, with the former merely constituting a reparameterization of the latter without enhancing representational capacity. Experiments demonstrate that the proposed method reduces GPU and CPU inference time by 44.3% and 58.1%, respectively, on speech enhancement tasks while maintaining performance, and consistently yields advantages across vision and natural language processing benchmarks.

0 citationsRead paper
Recent publications

Latest Papers

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

Jun 16, 2026

This work addresses the challenge of balancing model efficiency and performance in high-fidelity speech enhancement by proposing QC-GAN, a novel framework that integrates quaternion representations with the Conformer architecture for the first time. By leveraging Hamiltonian products to jointly model magnitude and phase in a structured manner, QC-GAN preserves their intrinsic correlation while substantially reducing parameter count. The approach further incorporates the MetricGAN training strategy and a metric learning-based discriminator to optimize perceptual quality. On the VoiceBank+DEMAND dataset, the model achieves a PESQ score of 3.48 with only 0.89 million parameters, and even a compact 35K-parameter variant attains 3.23—significantly outperforming conventional methods. Strong generalization capability is also demonstrated on the DNS-Challenge 3 benchmark.

0 citationsRead paper

Quaternion Self-Attention with Shared Scores

May 24, 2026

Existing quaternion self-attention mechanisms compute attention scores independently for each quaternion component, resulting in high computational overhead and inconsistent attention distributions. This work proposes a shared-score quaternion self-attention mechanism that generates a single real-valued attention score via quaternion inner product and shares the resulting attention distribution across all components. Theoretical analysis reveals that when queries and keys are pre-mixed through quaternion linear projections, component-wise independent scoring and shared scoring operate within the same interaction subspace, with the former merely constituting a reparameterization of the latter without enhancing representational capacity. Experiments demonstrate that the proposed method reduces GPU and CPU inference time by 44.3% and 58.1%, respectively, on speech enhancement tasks while maintaining performance, and consistently yields advantages across vision and natural language processing benchmarks.

0 citationsRead paper