Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning
研究通过约束训练目标与基础模型的知识对齐来减少事实性幻觉,提出两种新方法:证据重写和回忆重写,实验表明这些方法有效。
研究通过约束训练目标与基础模型的知识对齐来减少事实性幻觉,提出两种新方法:证据重写和回忆重写,实验表明这些方法有效。
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
This study investigates whether the well-established log-linear relationship between language model perplexity (PPL) and word error rate (WER) in automatic speech recognition (ASR) still holds for modern end-to-end ASR systems. Given that contemporary architectures commonly incorporate internal language modeling (ILM) capabilities, we systematically evaluate the efficacy of external language models, examining the PPL–WER correlation, the impact of encoder context length, and the applicability of large language models. For the first time, we comprehensively analyze how ILM—and its removal—affects the PPL–WER relationship. Our experiments reveal that while external language models continue to yield performance gains, the presence of ILM substantially shifts the PPL–WER trend; this shift is markedly altered upon ILM subtraction, indicating that neglecting internal language modeling leads to misleading assessments of external language model effectiveness.
This work addresses the limitation of BEST-RQ’s fixed online quantization mechanism, which constrains pseudo-label quality and yields weak supervision signals in self-supervised speech representation learning. While preserving the model’s architectural simplicity, the authors propose three key enhancements: replacing the linear projection with principal component analysis (PCA), iterative codebook optimization, and distillation-assisted codebook updates. These modifications substantially improve pseudo-label fidelity. Experimental results on the LibriSpeech dataset demonstrate that the proposed approach reduces the word error rate on the test-other subset from 10.1% to 8.8%, achieving a relative improvement of 12%.
This work addresses a critical yet previously unreported issue in memristor-based analog computing: positional encoding induces an excessively large output dynamic range, leading to severe distortion during analog-to-digital conversion (ADC). To mitigate this problem, the study introduces a hardware-aware optimization strategy for positional encoding. In scenarios with tunable ADCs, the approach jointly adjusts the weight scaling of the memristor layer and the ADC bit precision; when the ADC is fixed, it eliminates the linear transformation associated with positional encoding. Notably, the proposed method incurs no additional energy overhead and reduces computational degradation by approximately 50% and 30% in the respective scenarios, substantially improving analog computing accuracy for speech recognition tasks.
研究通过约束训练目标与基础模型的知识对齐来减少事实性幻觉,提出两种新方法:证据重写和回忆重写,实验表明这些方法有效。
This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.
This study investigates whether the well-established log-linear relationship between language model perplexity (PPL) and word error rate (WER) in automatic speech recognition (ASR) still holds for modern end-to-end ASR systems. Given that contemporary architectures commonly incorporate internal language modeling (ILM) capabilities, we systematically evaluate the efficacy of external language models, examining the PPL–WER correlation, the impact of encoder context length, and the applicability of large language models. For the first time, we comprehensively analyze how ILM—and its removal—affects the PPL–WER relationship. Our experiments reveal that while external language models continue to yield performance gains, the presence of ILM substantially shifts the PPL–WER trend; this shift is markedly altered upon ILM subtraction, indicating that neglecting internal language modeling leads to misleading assessments of external language model effectiveness.
This work addresses the limitation of BEST-RQ’s fixed online quantization mechanism, which constrains pseudo-label quality and yields weak supervision signals in self-supervised speech representation learning. While preserving the model’s architectural simplicity, the authors propose three key enhancements: replacing the linear projection with principal component analysis (PCA), iterative codebook optimization, and distillation-assisted codebook updates. These modifications substantially improve pseudo-label fidelity. Experimental results on the LibriSpeech dataset demonstrate that the proposed approach reduces the word error rate on the test-other subset from 10.1% to 8.8%, achieving a relative improvement of 12%.
This work addresses a critical yet previously unreported issue in memristor-based analog computing: positional encoding induces an excessively large output dynamic range, leading to severe distortion during analog-to-digital conversion (ADC). To mitigate this problem, the study introduces a hardware-aware optimization strategy for positional encoding. In scenarios with tunable ADCs, the approach jointly adjusts the weight scaling of the memristor layer and the ADC bit precision; when the ADC is fixed, it eliminates the linear transformation associated with positional encoding. Notably, the proposed method incurs no additional energy overhead and reduces computational degradation by approximately 50% and 30% in the respective scenarios, substantially improving analog computing accuracy for speech recognition tasks.