Institution profile

Huiyan Technology (Tianjin) Co., Ltd

Industry researchasia · cn
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection

Jun 30, 2026

This study addresses the limitations of existing Alzheimer’s disease (AD) speech-based detection methods, which often overlook the disruption of nonlinear linguistic structures and clinical heterogeneity. To this end, the authors propose a “Content–Structure–Flow” multi-view graph representation framework that leverages automatic speech recognition to construct semantic, dependency, and pointwise mutual information (PMI) co-occurrence graphs, effectively capturing narrative logic deviations. An heterogeneity-aware adaptive gating mechanism is further introduced to dynamically fuse these multi-view graphs, enhancing robustness across diverse populations. Integrating graph attention networks with the proposed multi-view fusion strategy, the model achieves a classification accuracy of 90.00% on the ADReSSo dataset. Ablation studies confirm the effectiveness and necessity of each component in the framework.

0 citationsRead paper

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

May 10, 2026

This work addresses the susceptibility of existing audio-visual large language models to cross-modal interference during inference, which often leads to hallucinations. To mitigate this issue, the authors propose a “separate-then-fuse” framework that first performs modality-specific chain-of-thought reasoning independently on audio and visual inputs, then fuses the resulting evidence to generate answers. Additionally, they introduce modality preference labels as auxiliary rewards in reinforcement learning, coupled with a data-driven preference annotation pipeline to model instance-level modality preferences. Evaluated on standard audio-visual question answering (AVQA) benchmarks, the method achieves a relative accuracy improvement of 5.16% on general tasks and 11.17% on cross-modal hallucination benchmarks, substantially enhancing model robustness and reliability.

0 citationsRead paper

MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates

Mar 24, 2026

This work addresses the limitation of existing self-supervised speech learning methods, which typically support only a single sampling rate and suffer performance degradation when trained on mixed-rate data due to temporal resolution mismatches. To overcome this, we propose MSR-HuBERT, the first framework enabling multi-sampling-rate self-supervised pretraining without resampling. MSR-HuBERT introduces a multi-sampling-rate adaptive downsampling CNN that maps raw waveforms of varying sampling rates—ranging from 16 kHz to 48 kHz—to a unified time resolution while preserving their original structure, thereby maintaining compatibility with HuBERT’s masked prediction objective and Transformer encoder. Experiments demonstrate that MSR-HuBERT outperforms standard HuBERT in both automatic speech recognition and full-band speech reconstruction tasks, effectively retaining high-frequency details and low-frequency semantic structures.

0 citationsRead paper

Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer's Disease Detection via Speech

Feb 16, 2026

This study addresses the challenge of data inefficiency in Alzheimer’s disease (AD) detection from speech, which stems from the scarcity of medical data and stringent privacy constraints. To overcome these limitations, the authors propose a novel framework that integrates speech content recombination-based augmentation, adaptive federated learning, and an attention-driven cross-modal alignment mechanism between acoustic and textual representations. This approach enables privacy-preserving collaboration across institutions while significantly enhancing data efficiency. The method achieves breakthroughs in absolute data efficiency, collaborative training efficiency, and representation learning efficiency, attaining a multimodal accuracy of 91.52% on the ADReSSo dataset—substantially outperforming existing centralized baselines.

0 citationsRead paper

POTSA: A Cross-Lingual Speech Alignment Framework for Low Resource Speech-to-Text Translation

Nov 12, 2025

Existing cross-lingual speech-to-text translation (S2TT) methods neglect semantic commonalities across source languages, limiting performance in low-resource and zero-shot settings. To address this, we propose POTSA—a Parallel Optimal Transport-based cross-lingual speech alignment framework—introducing optimal transport (OT) to low-resource S2TT for the first time. POTSA employs a Q-Former-driven token-level OT constraint, a bias compensation module, and a layer-wise scheduling strategy to progressively align cross-lingual speech representations from coarse- to fine-grained levels. Evaluated on the FLEURS benchmark, POTSA achieves state-of-the-art performance using only 10 hours of parallel speech per language: it improves average BLEU by 0.93 points across five high-resource languages and by 5.05 points on zero-shot languages. These gains demonstrate significantly enhanced multilingual semantic consistency and generalization capability.

0 citationsRead paper
Recent publications

Latest Papers

Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection

Jun 30, 2026

This study addresses the limitations of existing Alzheimer’s disease (AD) speech-based detection methods, which often overlook the disruption of nonlinear linguistic structures and clinical heterogeneity. To this end, the authors propose a “Content–Structure–Flow” multi-view graph representation framework that leverages automatic speech recognition to construct semantic, dependency, and pointwise mutual information (PMI) co-occurrence graphs, effectively capturing narrative logic deviations. An heterogeneity-aware adaptive gating mechanism is further introduced to dynamically fuse these multi-view graphs, enhancing robustness across diverse populations. Integrating graph attention networks with the proposed multi-view fusion strategy, the model achieves a classification accuracy of 90.00% on the ADReSSo dataset. Ablation studies confirm the effectiveness and necessity of each component in the framework.

0 citationsRead paper

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

May 10, 2026

This work addresses the susceptibility of existing audio-visual large language models to cross-modal interference during inference, which often leads to hallucinations. To mitigate this issue, the authors propose a “separate-then-fuse” framework that first performs modality-specific chain-of-thought reasoning independently on audio and visual inputs, then fuses the resulting evidence to generate answers. Additionally, they introduce modality preference labels as auxiliary rewards in reinforcement learning, coupled with a data-driven preference annotation pipeline to model instance-level modality preferences. Evaluated on standard audio-visual question answering (AVQA) benchmarks, the method achieves a relative accuracy improvement of 5.16% on general tasks and 11.17% on cross-modal hallucination benchmarks, substantially enhancing model robustness and reliability.

0 citationsRead paper

MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates

Mar 24, 2026

This work addresses the limitation of existing self-supervised speech learning methods, which typically support only a single sampling rate and suffer performance degradation when trained on mixed-rate data due to temporal resolution mismatches. To overcome this, we propose MSR-HuBERT, the first framework enabling multi-sampling-rate self-supervised pretraining without resampling. MSR-HuBERT introduces a multi-sampling-rate adaptive downsampling CNN that maps raw waveforms of varying sampling rates—ranging from 16 kHz to 48 kHz—to a unified time resolution while preserving their original structure, thereby maintaining compatibility with HuBERT’s masked prediction objective and Transformer encoder. Experiments demonstrate that MSR-HuBERT outperforms standard HuBERT in both automatic speech recognition and full-band speech reconstruction tasks, effectively retaining high-frequency details and low-frequency semantic structures.

0 citationsRead paper

Breaking Data Efficiency Dilemma: A Federated and Augmented Learning Framework For Alzheimer's Disease Detection via Speech

Feb 16, 2026

This study addresses the challenge of data inefficiency in Alzheimer’s disease (AD) detection from speech, which stems from the scarcity of medical data and stringent privacy constraints. To overcome these limitations, the authors propose a novel framework that integrates speech content recombination-based augmentation, adaptive federated learning, and an attention-driven cross-modal alignment mechanism between acoustic and textual representations. This approach enables privacy-preserving collaboration across institutions while significantly enhancing data efficiency. The method achieves breakthroughs in absolute data efficiency, collaborative training efficiency, and representation learning efficiency, attaining a multimodal accuracy of 91.52% on the ADReSSo dataset—substantially outperforming existing centralized baselines.

0 citationsRead paper

POTSA: A Cross-Lingual Speech Alignment Framework for Low Resource Speech-to-Text Translation

Nov 12, 2025

Existing cross-lingual speech-to-text translation (S2TT) methods neglect semantic commonalities across source languages, limiting performance in low-resource and zero-shot settings. To address this, we propose POTSA—a Parallel Optimal Transport-based cross-lingual speech alignment framework—introducing optimal transport (OT) to low-resource S2TT for the first time. POTSA employs a Q-Former-driven token-level OT constraint, a bias compensation module, and a layer-wise scheduling strategy to progressively align cross-lingual speech representations from coarse- to fine-grained levels. Evaluated on the FLEURS benchmark, POTSA achieves state-of-the-art performance using only 10 hours of parallel speech per language: it improves average BLEU by 0.93 points across five high-resource languages and by 5.05 points on zero-shot languages. These gains demonstrate significantly enhanced multilingual semantic consistency and generalization capability.

0 citationsRead paper