Orukeet: Multilingual ASR with Frozen Gabor Kernels

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过在Parakeet编码器中替换并冻结部分Gabor核,并使用多语言和多方言数据训练其余参数,以降低自动语音识别中的词错误率。
📝 Abstract
Orukeet replaces half of an adapted Parakeet encoder's temporal filters with 12,288 fitted Gabor kernels, freezes these replacements, and trains the remaining parameters on multilingual and multi-accent data. Final adaptation and checkpoint selection use LibriSpeech test-other. Across 20,146 FLEURS recordings in 25 languages, pooled word error rate (WER) falls from Parakeet's 11.01% to Orukeet's 9.85%, a 10.6% relative reduction. Orukeet has lower WER on 23 of the 25 languages. Orukeet outperforms Parakeet on 61 out of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). All comparisons decode the same audio with matched NeMo settings. The fitted kernels are stored as ordinary convolution weights, retaining Parakeet's architecture and inference operators.
Problem

Research questions and friction points this paper is trying to address.

multilingual ASR
word error rate
frozen Gabor kernels
temporal filters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gabor Kernels
Frozen Parameters
Multilingual ASR
Word Error Rate (WER) Reduction
🔎 Similar Papers
No similar papers found.