Improving Language Identification for Code-Switched Utterances with Integer Linear Programming
本文通过改进MaskLID算法并引入整数线性规划,解决了代码转换语句在语言识别中的不足,提高了多语言组合识别的准确性和可解释性。
本文通过改进MaskLID算法并引入整数线性规划,解决了代码转换语句在语言识别中的不足,提高了多语言组合识别的准确性和可解释性。
This study addresses the poor performance and inconsistent output in Central Kurdish speech translation, primarily caused by the scarcity of high-quality parallel data and the lack of orthographic standardization. To tackle these challenges, we present KUTED, the first large-scale English–Central Kurdish speech translation corpus comprising 91,000 sentence pairs, along with a systematic text normalization pipeline to unify orthographic conventions. Leveraging this resource, we train a from-scratch Transformer model and evaluate both the Seamless end-to-end system and a cascaded architecture combining NLLB and ASR components. Experimental results demonstrate that orthographic normalization yields a 3.0 BLEU improvement for Seamless on the FLEURS benchmark and achieves a BLEU score of 15.18 on an independent test set, significantly advancing speech translation for this low-resource language.
In low-resource neural machine translation (NMT), parallel bilingual corpora are scarce, yet monolingual target-language data are often abundant. Method: This paper proposes the first cross-lingual retrieval-augmented NMT framework explicitly leveraging monolingual target-language corpora—bypassing conventional bilingual memory banks. It retrieves relevant target-language segments directly using the source sentence as a query and introduces a novel joint sentence- and word-level contrastive learning objective to enforce multi-granularity semantic alignment, seamlessly integrated into the RANMT architecture. Contribution/Results: Experiments demonstrate that our approach significantly outperforms baseline NMT models and generic cross-lingual retrievers in both controlled settings and realistic low-resource scenarios. Notably, when the scale of target monolingual data vastly exceeds that of available parallel data, the method yields substantial BLEU improvements, validating its effectiveness in data-imbalanced low-resource regimes.
本文通过改进MaskLID算法并引入整数线性规划,解决了代码转换语句在语言识别中的不足,提高了多语言组合识别的准确性和可解释性。
This study addresses the poor performance and inconsistent output in Central Kurdish speech translation, primarily caused by the scarcity of high-quality parallel data and the lack of orthographic standardization. To tackle these challenges, we present KUTED, the first large-scale English–Central Kurdish speech translation corpus comprising 91,000 sentence pairs, along with a systematic text normalization pipeline to unify orthographic conventions. Leveraging this resource, we train a from-scratch Transformer model and evaluate both the Seamless end-to-end system and a cascaded architecture combining NLLB and ASR components. Experimental results demonstrate that orthographic normalization yields a 3.0 BLEU improvement for Seamless on the FLEURS benchmark and achieves a BLEU score of 15.18 on an independent test set, significantly advancing speech translation for this low-resource language.
In low-resource neural machine translation (NMT), parallel bilingual corpora are scarce, yet monolingual target-language data are often abundant. Method: This paper proposes the first cross-lingual retrieval-augmented NMT framework explicitly leveraging monolingual target-language corpora—bypassing conventional bilingual memory banks. It retrieves relevant target-language segments directly using the source sentence as a query and introduces a novel joint sentence- and word-level contrastive learning objective to enforce multi-granularity semantic alignment, seamlessly integrated into the RANMT architecture. Contribution/Results: Experiments demonstrate that our approach significantly outperforms baseline NMT models and generic cross-lingual retrievers in both controlled settings and realistic low-resource scenarios. Notably, when the scale of target monolingual data vastly exceeds that of available parallel data, the method yields substantial BLEU improvements, validating its effectiveness in data-imbalanced low-resource regimes.