Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过提出Hybrid Search方法,利用基础大语言模型与语音识别模型隐藏状态之间的互动特征来纠正和改进语音识别性能。
📝 Abstract
Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two strategies remains underexplored. In this work, we refine warm-initialized LLM-based ASR models by leveraging their own pre-adaptation base LLMs, focusing on LoRA-adapted settings where the base LLM is preserved. To achieve this, we propose Hybrid Search, a targeted correction strategy motivated by two observations. First, interaction features that characterize the relationship between LLM-based ASR hidden states and base-LLM hidden states provide informative signals about a token's degree of semantic dependence. Second, selectively refining targeted tokens with high semantic dependence improves ASR performance far beyond naive global LLM-correction methods including rescoring and late fusion. Our analysis suggests that, even after semantic knowledge transfer through warm initialization, LLM-based ASR models can still leverage their base LLM to further improve inference-time performance.
Problem

Research questions and friction points this paper is trying to address.

automatic speech recognition
large language models
warm initialization
logit fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Search
hidden-state interactions
semantic dependence
LoRA-adapted settings
targeted correction
🔎 Similar Papers
No similar papers found.