Institution profile

ASUS

Industry researchasia · tw
Official website
Research library25linked papers
Opportunities0open roles
Selected work

Representative Papers

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Jul 13, 2026

This work addresses the limited capability of existing large audio language models in perceiving fine-grained non-semantic acoustic attributes, such as vocal emotion. The authors propose a training-free, label-free inference-time intervention method that identifies and amplifies a small subset of neurons most sensitive to acoustic information by comparing their activation patterns in response to real speech versus noise reference signals. This approach enables, for the first time, neuron-level precise intervention within the audio encoder itself, revealing the critical roles of encoder depth and neuronal selectivity in acoustic perception. Experimental results demonstrate substantial improvements, with average accuracy gains of 25.7, 21.4, and 9.7 percentage points on Audio-Flamingo-3, Qwen2.5-Omni, and Kimi-Audio, respectively, significantly outperforming intervention strategies applied at the decoder or within the language model.

0 citationsRead paper

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

Jul 07, 2026

This work addresses the vector collapse problem in audio language models, where the use of a Q-Former connector to compress speech encoder outputs often causes representations to align along a single direction, thereby discarding paralinguistic information such as speaker identity, gender, and prosody. To mitigate this issue, the authors propose ORCA, the first method to introduce an inter-group orthogonality constraint: queries are partitioned into multiple groups, and their outputs are enforced to point in distinct directions, effectively restoring paralinguistic diversity. Built upon a Querying Transformer, ORCA employs a grouped orthogonal connector that substantially alleviates output collapse. Experiments demonstrate that ORCA achieves 75.2% accuracy on the SAKURA multi-hop reasoning task—surpassing the 4B baseline by 26.4 percentage points—while reducing query redundancy in the connector by 12× and increasing cross-speaker representation variance by 75×.

0 citationsRead paper

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

Jul 06, 2026

This work addresses the limited capacity of end-to-end automatic speech recognition (ASR) models to iteratively refine predictions for challenging inputs. The authors propose LatentASR, a method that enables efficient and stable test-time adaptation through continuous latent variables while keeping the backbone model parameters frozen. Requiring only approximately 4M additional parameters and a small activation set of 500 utterances, LatentASR employs a latent adapter for dynamic updates and integrates a value head to implement an input-dependent early stopping strategy. The approach achieves consistent performance gains across diverse benchmarks: relative word error rate (WER) reductions of 2.54% and 0.47% on FLEURS and VoxPopuli, respectively, and a 16.0% relative character error rate (CER) improvement on ASCEND data featuring accented speech and code-switching, with consistent enhancements observed across 30 languages.

0 citationsRead paper

Context-Aware ASR for Mandarin Technical Lectures

Jul 06, 2026

This study addresses the challenge of evaluating automatic speech recognition (ASR) performance on Mandarin technical lectures, where critical English terms are frequently interspersed and conventional character error rate (CER) metrics fail to capture term-level accuracy. To overcome this limitation, the authors propose a reference-free, two-stage decoding approach: the first stage performs standard ASR and automatically extracts high-frequency technical terms from the output to construct a dynamic, context-aware lexicon; the second stage leverages this lexicon to guide rescoring and refine recognition results. This work introduces a novel term-centric evaluation metric that exposes blind spots in CER and demonstrates consistent improvements across five mainstream ASR systems. Experimental results show substantial gains in term recall (up to 62.05% with Breeze-ASR-25) and precision (82.73%), while maintaining or even reducing overall CER.

0 citationsRead paper
Recent publications

Latest Papers

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Jul 13, 2026

This work addresses the limited capability of existing large audio language models in perceiving fine-grained non-semantic acoustic attributes, such as vocal emotion. The authors propose a training-free, label-free inference-time intervention method that identifies and amplifies a small subset of neurons most sensitive to acoustic information by comparing their activation patterns in response to real speech versus noise reference signals. This approach enables, for the first time, neuron-level precise intervention within the audio encoder itself, revealing the critical roles of encoder depth and neuronal selectivity in acoustic perception. Experimental results demonstrate substantial improvements, with average accuracy gains of 25.7, 21.4, and 9.7 percentage points on Audio-Flamingo-3, Qwen2.5-Omni, and Kimi-Audio, respectively, significantly outperforming intervention strategies applied at the decoder or within the language model.

0 citationsRead paper

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

Jul 07, 2026

This work addresses the vector collapse problem in audio language models, where the use of a Q-Former connector to compress speech encoder outputs often causes representations to align along a single direction, thereby discarding paralinguistic information such as speaker identity, gender, and prosody. To mitigate this issue, the authors propose ORCA, the first method to introduce an inter-group orthogonality constraint: queries are partitioned into multiple groups, and their outputs are enforced to point in distinct directions, effectively restoring paralinguistic diversity. Built upon a Querying Transformer, ORCA employs a grouped orthogonal connector that substantially alleviates output collapse. Experiments demonstrate that ORCA achieves 75.2% accuracy on the SAKURA multi-hop reasoning task—surpassing the 4B baseline by 26.4 percentage points—while reducing query redundancy in the connector by 12× and increasing cross-speaker representation variance by 75×.

0 citationsRead paper

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

Jul 06, 2026

This work addresses the limited capacity of end-to-end automatic speech recognition (ASR) models to iteratively refine predictions for challenging inputs. The authors propose LatentASR, a method that enables efficient and stable test-time adaptation through continuous latent variables while keeping the backbone model parameters frozen. Requiring only approximately 4M additional parameters and a small activation set of 500 utterances, LatentASR employs a latent adapter for dynamic updates and integrates a value head to implement an input-dependent early stopping strategy. The approach achieves consistent performance gains across diverse benchmarks: relative word error rate (WER) reductions of 2.54% and 0.47% on FLEURS and VoxPopuli, respectively, and a 16.0% relative character error rate (CER) improvement on ASCEND data featuring accented speech and code-switching, with consistent enhancements observed across 30 languages.

0 citationsRead paper

Context-Aware ASR for Mandarin Technical Lectures

Jul 06, 2026

This study addresses the challenge of evaluating automatic speech recognition (ASR) performance on Mandarin technical lectures, where critical English terms are frequently interspersed and conventional character error rate (CER) metrics fail to capture term-level accuracy. To overcome this limitation, the authors propose a reference-free, two-stage decoding approach: the first stage performs standard ASR and automatically extracts high-frequency technical terms from the output to construct a dynamic, context-aware lexicon; the second stage leverages this lexicon to guide rescoring and refine recognition results. This work introduces a novel term-centric evaluation metric that exposes blind spots in CER and demonstrates consistent improvements across five mainstream ASR systems. Experimental results show substantial gains in term recall (up to 62.05% with Breeze-ASR-25) and precision (82.73%), while maintaining or even reducing overall CER.

0 citationsRead paper