Institution profile

aiOla

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Drax: Speech Recognition with Discrete Flow Matching

Oct 05, 2025

To address the performance bottleneck in non-autoregressive (NAR) automatic speech recognition (ASR) caused by train-inference distribution mismatch, this paper proposes Drax—the first NAR ASR framework based on discrete flow matching. Drax constructs an audio-conditioned probability flow path that explicitly models intermediate erroneous token trajectories during inference, thereby mitigating distributional shift between training and inference. Theoretically, it establishes a connection between generalization error and cumulative velocity error, providing principled guidance for model design. Drax enables fully parallel decoding and achieves recognition accuracy competitive with state-of-the-art autoregressive models on benchmarks including LibriSpeech, while substantially improving decoding efficiency. Extensive experiments validate the effectiveness and scalability of discrete flow matching for ASR tasks.

0 citationsRead paper

Beyond Transcription: Mechanistic Interpretability in ASR

Aug 21, 2025

Current ASR systems suffer from a fundamental lack of mechanistic understanding—particularly regarding how acoustic and semantic information dynamically evolve across layers—resulting in severe interpretability deficits. This work pioneers the systematic application of mechanistic interpretability techniques—including logit lens analysis, linear probing, and activation patching—to encoder-decoder ASR models, enabling layer-wise dissection of representational dynamics. We identify critical cross-layer interaction pathways responsible for repetition hallucinations and, for the first time, uncover latent semantic bias in deep acoustic representations: speech features undergo premature and excessive semantic grounding. These findings elucidate previously unknown internal mechanisms of ASR models and establish novel theoretical foundations—along with actionable intervention points—for enhancing model transparency, robustness, and controllability.

0 citationsRead paper

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Jun 11, 2025

Current text-to-speech (TTS) models generate natural speech but struggle with text-driven, controllable co-synthesis of speech and complex environmental sounds due to the absence of real-world aligned speech–environment audio pairs. To address this, we propose the first end-to-end environment-aware TTS framework based on flow matching for joint generation. We introduce a novel self-supervised acoustic disentanglement method that decomposes unlabeled recordings into speech, text, and background components—overcoming the data scarcity bottleneck. Furthermore, we design a joint conditional modeling scheme with fine-grained environmental intensity control. Experiments demonstrate significant improvements over state-of-the-art models across objective and subjective metrics—including naturalness, environmental consistency, and scene diversity—achieving, for the first time, high-fidelity, controllable, text-driven joint synthesis of speech and environmental audio.

0 citationsRead paper

FlowTSE: Target Speaker Extraction with Flow Matching

May 20, 2025

Target Speaker Extraction (TSE) aims to isolate a target speaker’s speech from a mixture using a registered enrollment utterance; however, existing generative approaches often rely on complex pipelines and pre-trained components, resulting in high computational overhead and limited modeling capacity. This paper proposes FlowTSE—the first end-to-end TSE framework leveraging Conditional Flow Matching (CFM), which directly generates the target speaker’s spectrogram conditioned on the mixture’s mel-spectrogram and the enrollment speech. We further introduce a novel complex-valued STFT-conditioned vocoder, significantly improving phase reconstruction fidelity. FlowTSE requires no pre-trained modules or cascaded processing stages. Evaluated on standard TSE benchmarks, it matches or surpasses state-of-the-art baselines while offering superior performance, architectural simplicity, and computational efficiency.

0 citationsRead paper
Recent publications

Latest Papers

Drax: Speech Recognition with Discrete Flow Matching

Oct 05, 2025

To address the performance bottleneck in non-autoregressive (NAR) automatic speech recognition (ASR) caused by train-inference distribution mismatch, this paper proposes Drax—the first NAR ASR framework based on discrete flow matching. Drax constructs an audio-conditioned probability flow path that explicitly models intermediate erroneous token trajectories during inference, thereby mitigating distributional shift between training and inference. Theoretically, it establishes a connection between generalization error and cumulative velocity error, providing principled guidance for model design. Drax enables fully parallel decoding and achieves recognition accuracy competitive with state-of-the-art autoregressive models on benchmarks including LibriSpeech, while substantially improving decoding efficiency. Extensive experiments validate the effectiveness and scalability of discrete flow matching for ASR tasks.

0 citationsRead paper

Beyond Transcription: Mechanistic Interpretability in ASR

Aug 21, 2025

Current ASR systems suffer from a fundamental lack of mechanistic understanding—particularly regarding how acoustic and semantic information dynamically evolve across layers—resulting in severe interpretability deficits. This work pioneers the systematic application of mechanistic interpretability techniques—including logit lens analysis, linear probing, and activation patching—to encoder-decoder ASR models, enabling layer-wise dissection of representational dynamics. We identify critical cross-layer interaction pathways responsible for repetition hallucinations and, for the first time, uncover latent semantic bias in deep acoustic representations: speech features undergo premature and excessive semantic grounding. These findings elucidate previously unknown internal mechanisms of ASR models and establish novel theoretical foundations—along with actionable intervention points—for enhancing model transparency, robustness, and controllability.

0 citationsRead paper

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Jun 11, 2025

Current text-to-speech (TTS) models generate natural speech but struggle with text-driven, controllable co-synthesis of speech and complex environmental sounds due to the absence of real-world aligned speech–environment audio pairs. To address this, we propose the first end-to-end environment-aware TTS framework based on flow matching for joint generation. We introduce a novel self-supervised acoustic disentanglement method that decomposes unlabeled recordings into speech, text, and background components—overcoming the data scarcity bottleneck. Furthermore, we design a joint conditional modeling scheme with fine-grained environmental intensity control. Experiments demonstrate significant improvements over state-of-the-art models across objective and subjective metrics—including naturalness, environmental consistency, and scene diversity—achieving, for the first time, high-fidelity, controllable, text-driven joint synthesis of speech and environmental audio.

0 citationsRead paper

FlowTSE: Target Speaker Extraction with Flow Matching

May 20, 2025

Target Speaker Extraction (TSE) aims to isolate a target speaker’s speech from a mixture using a registered enrollment utterance; however, existing generative approaches often rely on complex pipelines and pre-trained components, resulting in high computational overhead and limited modeling capacity. This paper proposes FlowTSE—the first end-to-end TSE framework leveraging Conditional Flow Matching (CFM), which directly generates the target speaker’s spectrogram conditioned on the mixture’s mel-spectrogram and the enrollment speech. We further introduce a novel complex-valued STFT-conditioned vocoder, significantly improving phase reconstruction fidelity. FlowTSE requires no pre-trained modules or cascaded processing stages. Evaluated on standard TSE benchmarks, it matches or surpasses state-of-the-art baselines while offering superior performance, architectural simplicity, and computational efficiency.

0 citationsRead paper