Institution profile

AppTek GmbH

Industry researcheurope · de
Official website
Research library23linked papers
Opportunities0open roles
Selected work

Representative Papers

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

Jul 07, 2026

This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.

0 citationsRead paper

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

Jul 06, 2026

This study investigates whether the well-established log-linear relationship between language model perplexity (PPL) and word error rate (WER) in automatic speech recognition (ASR) still holds for modern end-to-end ASR systems. Given that contemporary architectures commonly incorporate internal language modeling (ILM) capabilities, we systematically evaluate the efficacy of external language models, examining the PPL–WER correlation, the impact of encoder context length, and the applicability of large language models. For the first time, we comprehensively analyze how ILM—and its removal—affects the PPL–WER relationship. Our experiments reveal that while external language models continue to yield performance gains, the presence of ILM substantially shifts the PPL–WER trend; this shift is markedly altered upon ILM subtraction, indicating that neglecting internal language modeling leads to misleading assessments of external language model effectiveness.

0 citationsRead paper

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

Jun 24, 2026

This work addresses the limitation of BEST-RQ’s fixed online quantization mechanism, which constrains pseudo-label quality and yields weak supervision signals in self-supervised speech representation learning. While preserving the model’s architectural simplicity, the authors propose three key enhancements: replacing the linear projection with principal component analysis (PCA), iterative codebook optimization, and distillation-assisted codebook updates. These modifications substantially improve pseudo-label fidelity. Experimental results on the LibriSpeech dataset demonstrate that the proposed approach reduces the word error rate on the test-other subset from 10.1% to 8.8%, achieving a relative improvement of 12%.

0 citationsRead paper

Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition

Jun 11, 2026

This work addresses a critical yet previously unreported issue in memristor-based analog computing: positional encoding induces an excessively large output dynamic range, leading to severe distortion during analog-to-digital conversion (ADC). To mitigate this problem, the study introduces a hardware-aware optimization strategy for positional encoding. In scenarios with tunable ADCs, the approach jointly adjusts the weight scaling of the memristor layer and the ADC bit precision; when the ADC is fixed, it eliminates the linear transformation associated with positional encoding. Notably, the proposed method incurs no additional energy overhead and reduces computational degradation by approximately 50% and 30% in the respective scenarios, substantially improving analog computing accuracy for speech recognition tasks.

0 citationsRead paper
Recent publications

Latest Papers

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

Jul 07, 2026

This work addresses the limited word-level time alignment capability of current automatic speech recognition (ASR) models—such as attention-based encoder-decoder (AED) systems and speech large language models—which often lack precise temporal grounding, while conventional alignment methods are constrained by encoder frame rates and offer only modest accuracy. The authors propose a general, training-free, and model-agnostic gradient-driven alignment approach that computes frame-level saliency maps via gradients of token log-probabilities with respect to the input signal under teacher forcing, followed by dynamic programming to decode word boundaries. Applicable to any differentiable ASR model, this method achieves high-precision alignment at the original input sampling rate. Experiments across 16 models on TIMIT and Buckeye datasets show that, although slightly less accurate than strong native aligners, it outperforms them in scenarios where native alignment capabilities are weak, such as with streaming ASR models.

0 citationsRead paper

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

Jul 06, 2026

This study investigates whether the well-established log-linear relationship between language model perplexity (PPL) and word error rate (WER) in automatic speech recognition (ASR) still holds for modern end-to-end ASR systems. Given that contemporary architectures commonly incorporate internal language modeling (ILM) capabilities, we systematically evaluate the efficacy of external language models, examining the PPL–WER correlation, the impact of encoder context length, and the applicability of large language models. For the first time, we comprehensively analyze how ILM—and its removal—affects the PPL–WER relationship. Our experiments reveal that while external language models continue to yield performance gains, the presence of ILM substantially shifts the PPL–WER trend; this shift is markedly altered upon ILM subtraction, indicating that neglecting internal language modeling leads to misleading assessments of external language model effectiveness.

0 citationsRead paper

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

Jun 24, 2026

This work addresses the limitation of BEST-RQ’s fixed online quantization mechanism, which constrains pseudo-label quality and yields weak supervision signals in self-supervised speech representation learning. While preserving the model’s architectural simplicity, the authors propose three key enhancements: replacing the linear projection with principal component analysis (PCA), iterative codebook optimization, and distillation-assisted codebook updates. These modifications substantially improve pseudo-label fidelity. Experimental results on the LibriSpeech dataset demonstrate that the proposed approach reduces the word error rate on the test-other subset from 10.1% to 8.8%, achieving a relative improvement of 12%.

0 citationsRead paper

Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition

Jun 11, 2026

This work addresses a critical yet previously unreported issue in memristor-based analog computing: positional encoding induces an excessively large output dynamic range, leading to severe distortion during analog-to-digital conversion (ADC). To mitigate this problem, the study introduces a hardware-aware optimization strategy for positional encoding. In scenarios with tunable ADCs, the approach jointly adjusts the weight scaling of the memristor layer and the ADC bit precision; when the ADC is fixed, it eliminates the linear transformation associated with positional encoding. Notably, the proposed method incurs no additional energy overhead and reduces computational degradation by approximately 50% and 30% in the respective scenarios, substantially improving analog computing accuracy for speech recognition tasks.

0 citationsRead paper