T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition
研究针对台语自动语音识别中变调问题,提出T-SANDHI模型,通过解耦表面声学和词典意图,利用多任务学习和动态门控方法有效解决音调混淆问题。
研究针对台语自动语音识别中变调问题,提出T-SANDHI模型,通过解耦表面声学和词典意图,利用多任务学习和动态门控方法有效解决音调混淆问题。
To address the challenges in Taiwanese Hokkien ASR—namely, the difficulty of modeling fine-grained phonetic details using character-based annotations and the limited lexical-syntactic coverage of romanized (e.g., Tâi-lô) transcriptions—this paper proposes a cross-lingual, two-stage fine-tuning framework. First, it leverages Tâi-lô romanization to train HuBERT for learning phoneme- and tone-aware acoustic representations. Second, it incorporates character-level text to jointly model lexical and syntactic structures, enabling synergistic alignment between acoustic and orthographic information. The method innovatively integrates dual annotation modalities, circumventing the limitations inherent to single-modality paradigms. Evaluated on the TAT-MOE benchmark, our approach achieves a 24.88% relative reduction in character error rate over strong baselines. The model is parameter-efficient and scalable, offering a reusable technical pathway for low-resource dialectal ASR.
研究针对台语自动语音识别中变调问题,提出T-SANDHI模型,通过解耦表面声学和词典意图,利用多任务学习和动态门控方法有效解决音调混淆问题。
To address the challenges in Taiwanese Hokkien ASR—namely, the difficulty of modeling fine-grained phonetic details using character-based annotations and the limited lexical-syntactic coverage of romanized (e.g., Tâi-lô) transcriptions—this paper proposes a cross-lingual, two-stage fine-tuning framework. First, it leverages Tâi-lô romanization to train HuBERT for learning phoneme- and tone-aware acoustic representations. Second, it incorporates character-level text to jointly model lexical and syntactic structures, enabling synergistic alignment between acoustic and orthographic information. The method innovatively integrates dual annotation modalities, circumventing the limitations inherent to single-modality paradigms. Evaluated on the TAT-MOE benchmark, our approach achieves a 24.88% relative reduction in character error rate over strong baselines. The model is parameter-efficient and scalable, offering a reusable technical pathway for low-resource dialectal ASR.