Institution profile

EZAI

Industry research
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition

Nov 10, 2025

To address the challenges in Taiwanese Hokkien ASR—namely, the difficulty of modeling fine-grained phonetic details using character-based annotations and the limited lexical-syntactic coverage of romanized (e.g., Tâi-lô) transcriptions—this paper proposes a cross-lingual, two-stage fine-tuning framework. First, it leverages Tâi-lô romanization to train HuBERT for learning phoneme- and tone-aware acoustic representations. Second, it incorporates character-level text to jointly model lexical and syntactic structures, enabling synergistic alignment between acoustic and orthographic information. The method innovatively integrates dual annotation modalities, circumventing the limitations inherent to single-modality paradigms. Evaluated on the TAT-MOE benchmark, our approach achieves a 24.88% relative reduction in character error rate over strong baselines. The model is parameter-efficient and scalable, offering a reusable technical pathway for low-resource dialectal ASR.

0 citationsRead paper
Recent publications

Latest Papers

CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition

Nov 10, 2025

To address the challenges in Taiwanese Hokkien ASR—namely, the difficulty of modeling fine-grained phonetic details using character-based annotations and the limited lexical-syntactic coverage of romanized (e.g., Tâi-lô) transcriptions—this paper proposes a cross-lingual, two-stage fine-tuning framework. First, it leverages Tâi-lô romanization to train HuBERT for learning phoneme- and tone-aware acoustic representations. Second, it incorporates character-level text to jointly model lexical and syntactic structures, enabling synergistic alignment between acoustic and orthographic information. The method innovatively integrates dual annotation modalities, circumventing the limitations inherent to single-modality paradigms. Evaluated on the TAT-MOE benchmark, our approach achieves a 24.88% relative reduction in character error rate over strong baselines. The model is parameter-efficient and scalable, offering a reusable technical pathway for low-resource dialectal ASR.

0 citationsRead paper